What this benchmark answers
This benchmark compares 30 models under the same OpenUI task: 46 interface briefs, four generations per brief, scored for structural validity. The chart pairs that score with measured list-price cost per generated task where a comparable API price exists.
A high score means the generated component graph parsed, produced a root, resolved its references, used valid required and enum props, avoided truncation, and met the shared component-count floor. It is not a general intelligence score.
All model results
All 30 rows are present below, including the five compact or local models deselected by default in the visual chart, which opens on models at or above 70%.
| Model | Provider | Valid | Cost / task | Pricing | Pareto | Default view |
|---|---|---|---|---|---|---|
| Grok 4.6 | xAI | 99.5% | $0.0185 | provider-list-price | Yes | Shown |
| GPT-5.6 Sol | OpenAI | 99.5% | $0.0476 | provider-list-price | No | Shown |
| Claude Opus 4.8 | Anthropic | 98.9% | $0.0493 | provider-list-price | No | Shown |
| Gemini 3.7 Flash | 98.9% | $0.0100 | provider-list-price | Yes | Shown | |
| Claude Sonnet 5 | Anthropic | 98.4% | $0.0202 | provider-list-price | No | Shown |
| GPT-5.6 Terra | OpenAI | 98.4% | $0.0241 | provider-list-price | No | Shown |
| Kimi K3 | Moonshot | 96.2% | $0.0333 | provider-list-price | No | Shown |
| Claude Opus 5 | Anthropic | 96.2% | $0.0774 | provider-list-price | No | Shown |
| Muse Spark 1.2 | Meta | 96.2% | $0.0157 | provider-list-price | No | Shown |
| Claude Sonnet 4.6 | Anthropic | 92.9% | $0.0450 | provider-list-price | No | Shown |
| Qwen3.8 2.4T | Alibaba | 91.8% | $0.0170 | provider-list-price | No | Shown |
| GLM-5.3 | Zhipu | 90.8% | $0.0120 | provider-list-price | No | Shown |
| Inkling Small | Thinking Machines | 88.6% | $0.0033 | provider-list-price | Yes | Shown |
| DeepSeek V4 Flash | DeepSeek | 85.8% | $0.0004 | provider-list-price | Yes | Shown |
| DeepSeek V4 Pro | DeepSeek | 84.2% | $0.0089 | provider-list-price | No | Shown |
| GPT-5.6 Luna | OpenAI | 83.7% | $0.0022 | provider-list-price | No | Shown |
| Qwen3.8 27B | Alibaba | 78.8% | $0.0054 | provider-list-price | No | Shown |
| Gemini 3.6 Flash | 78.8% | $0.0091 | provider-list-price | No | Shown | |
| Gemini 3.5 Flash-Lite | 78.3% | $0.0041 | provider-list-price | No | Shown | |
| Inkling | Thinking Machines | 73.4% | $0.0080 | provider-list-price | No | Shown |
| Qwen3.6 27B | Alibaba | 68.5% | $0.0063 | provider-list-price | No | Hidden |
| Qwen3.6 35B-A3B | Alibaba | 61.4% | $0.0017 | provider-list-price | No | Hidden |
| Gemma 4 31B | 46.7% | $0.0009 | provider-list-price | No | Hidden | |
| Phi-4 | Microsoft | 44.0% | $0.0004 | provider-list-price | No | Hidden |
| Gemma 4 26B-A4B | 29.9% | $0.0007 | provider-list-price | No | Hidden | |
| Ministral 8B | Mistral | 27.2% | $0.0009 | provider-list-price | No | Hidden |
| Granite 4.1 8B | IBM | 14.7% | $0.0002 | provider-list-price | Yes | Hidden |
| LFM 2.5 2.6B | Liquid | 3.3% | $0 / free | free | Yes | Hidden |
| DiffusionGemma 26B-A4B | 13.0% | Not comparable | self-hosted | No | Hidden | |
| Ling 3.0 Tiny | InclusionAI | 9.8% | Not comparable | self-hosted | No | Hidden |
How to interpret it
- Provider color identifies a company; it is not a quality highlight.
- A Pareto point is not dominated on both measured cost and structural validity. It is not a general model ranking.
- Self-hosted cost is recorded as null because there is no comparable API list price. Null does not mean free.
- Every row is scored by one build of the shipped OpenUI parser, lang-core 0.2.16, whose parser validates enum and scalar prop values.
The exact generation and scoring condition is documented on the methodology page.