# OpenUI language and model benchmark

Canonical page: https://www.openui.com/benchmarks/language

## What this benchmark answers

How reliably and economically do different models generate structurally valid OpenUI?

## Scope

- 30 models.
- 46 interface briefs and four generations per brief.
- Structural validity is the vertical metric; measured list-price cost per task is the horizontal metric where comparable.
- Self-hosted cost is null, not zero.
- Models below 70% structural validity are deselected by default in the visual chart but remain in this table.
- Every row is scored by one build of the shipped OpenUI parser, lang-core 0.2.16.

| Model | Provider | Structural validity | Cost per task | Cost type | Pareto frontier | Shown by default |
| --- | --- | ---: | ---: | --- | --- | --- |
| Grok 4.6 | xAI | 99.5% | $0.0185 | provider-list-price | yes | yes |
| GPT-5.6 Sol | OpenAI | 99.5% | $0.0476 | provider-list-price | no | yes |
| Claude Opus 4.8 | Anthropic | 98.9% | $0.0493 | provider-list-price | no | yes |
| Gemini 3.7 Flash | Google | 98.9% | $0.0100 | provider-list-price | yes | yes |
| Claude Sonnet 5 | Anthropic | 98.4% | $0.0202 | provider-list-price | no | yes |
| GPT-5.6 Terra | OpenAI | 98.4% | $0.0241 | provider-list-price | no | yes |
| Kimi K3 | Moonshot | 96.2% | $0.0333 | provider-list-price | no | yes |
| Claude Opus 5 | Anthropic | 96.2% | $0.0774 | provider-list-price | no | yes |
| Muse Spark 1.2 | Meta | 96.2% | $0.0157 | provider-list-price | no | yes |
| Claude Sonnet 4.6 | Anthropic | 92.9% | $0.0450 | provider-list-price | no | yes |
| Qwen3.8 2.4T | Alibaba | 91.8% | $0.0170 | provider-list-price | no | yes |
| GLM-5.3 | Zhipu | 90.8% | $0.0120 | provider-list-price | no | yes |
| Inkling Small | Thinking Machines | 88.6% | $0.0033 | provider-list-price | yes | yes |
| DeepSeek V4 Flash | DeepSeek | 85.8% | $0.0004 | provider-list-price | yes | yes |
| DeepSeek V4 Pro | DeepSeek | 84.2% | $0.0089 | provider-list-price | no | yes |
| GPT-5.6 Luna | OpenAI | 83.7% | $0.0022 | provider-list-price | no | yes |
| Qwen3.8 27B | Alibaba | 78.8% | $0.0054 | provider-list-price | no | yes |
| Gemini 3.6 Flash | Google | 78.8% | $0.0091 | provider-list-price | no | yes |
| Gemini 3.5 Flash-Lite | Google | 78.3% | $0.0041 | provider-list-price | no | yes |
| Inkling | Thinking Machines | 73.4% | $0.0080 | provider-list-price | no | yes |
| Qwen3.6 27B | Alibaba | 68.5% | $0.0063 | provider-list-price | no | no |
| Qwen3.6 35B-A3B | Alibaba | 61.4% | $0.0017 | provider-list-price | no | no |
| Gemma 4 31B | Google | 46.7% | $0.0009 | provider-list-price | no | no |
| Phi-4 | Microsoft | 44.0% | $0.0004 | provider-list-price | no | no |
| Gemma 4 26B-A4B | Google | 29.9% | $0.0007 | provider-list-price | no | no |
| Ministral 8B | Mistral | 27.2% | $0.0009 | provider-list-price | no | no |
| Granite 4.1 8B | IBM | 14.7% | $0.0002 | provider-list-price | yes | no |
| LFM 2.5 2.6B | Liquid | 3.3% | $0 / free | free | yes | no |
| DiffusionGemma 26B-A4B | Google | 13.0% | not comparable | self-hosted | no | no |
| Ling 3.0 Tiny | InclusionAI | 9.8% | not comparable | self-hosted | no | no |

## Links

- Methodology: https://www.openui.com/benchmarks/methodology
- JSON: https://www.openui.com/benchmarks/language/data.json
- CSV: https://www.openui.com/benchmarks/language/data.csv
- Combined benchmark: https://www.openui.com/benchmarks
