OpenUI benchmark · Models

OpenUI language and model benchmark

How reliably and economically do different models generate structurally valid OpenUI?

What this benchmark answers

This benchmark compares 30 models under the same OpenUI task: 46 interface briefs, four generations per brief, scored for structural validity. The chart pairs that score with measured list-price cost per generated task where a comparable API price exists.

A high score means the generated component graph parsed, produced a root, resolved its references, used valid required and enum props, avoided truncation, and met the shared component-count floor. It is not a general intelligence score.

All model results

All 30 rows are present below, including the five compact or local models deselected by default in the visual chart, which opens on models at or above 70%.

OpenUI structural validity and measured cost across models
ModelProviderValidCost / taskPricingParetoDefault view
Grok 4.6xAI99.5%$0.0185provider-list-priceYesShown
GPT-5.6 SolOpenAI99.5%$0.0476provider-list-priceNoShown
Claude Opus 4.8Anthropic98.9%$0.0493provider-list-priceNoShown
Gemini 3.7 FlashGoogle98.9%$0.0100provider-list-priceYesShown
Claude Sonnet 5Anthropic98.4%$0.0202provider-list-priceNoShown
GPT-5.6 TerraOpenAI98.4%$0.0241provider-list-priceNoShown
Kimi K3Moonshot96.2%$0.0333provider-list-priceNoShown
Claude Opus 5Anthropic96.2%$0.0774provider-list-priceNoShown
Muse Spark 1.2Meta96.2%$0.0157provider-list-priceNoShown
Claude Sonnet 4.6Anthropic92.9%$0.0450provider-list-priceNoShown
Qwen3.8 2.4TAlibaba91.8%$0.0170provider-list-priceNoShown
GLM-5.3Zhipu90.8%$0.0120provider-list-priceNoShown
Inkling SmallThinking Machines88.6%$0.0033provider-list-priceYesShown
DeepSeek V4 FlashDeepSeek85.8%$0.0004provider-list-priceYesShown
DeepSeek V4 ProDeepSeek84.2%$0.0089provider-list-priceNoShown
GPT-5.6 LunaOpenAI83.7%$0.0022provider-list-priceNoShown
Qwen3.8 27BAlibaba78.8%$0.0054provider-list-priceNoShown
Gemini 3.6 FlashGoogle78.8%$0.0091provider-list-priceNoShown
Gemini 3.5 Flash-LiteGoogle78.3%$0.0041provider-list-priceNoShown
InklingThinking Machines73.4%$0.0080provider-list-priceNoShown
Qwen3.6 27BAlibaba68.5%$0.0063provider-list-priceNoHidden
Qwen3.6 35B-A3BAlibaba61.4%$0.0017provider-list-priceNoHidden
Gemma 4 31BGoogle46.7%$0.0009provider-list-priceNoHidden
Phi-4Microsoft44.0%$0.0004provider-list-priceNoHidden
Gemma 4 26B-A4BGoogle29.9%$0.0007provider-list-priceNoHidden
Ministral 8BMistral27.2%$0.0009provider-list-priceNoHidden
Granite 4.1 8BIBM14.7%$0.0002provider-list-priceYesHidden
LFM 2.5 2.6BLiquid3.3%$0 / freefreeYesHidden
DiffusionGemma 26B-A4BGoogle13.0%Not comparableself-hostedNoHidden
Ling 3.0 TinyInclusionAI9.8%Not comparableself-hostedNoHidden

How to interpret it

  • Provider color identifies a company; it is not a quality highlight.
  • A Pareto point is not dominated on both measured cost and structural validity. It is not a general model ranking.
  • Self-hosted cost is recorded as null because there is no comparable API list price. Null does not mean free.
  • Every row is scored by one build of the shipped OpenUI parser, lang-core 0.2.16, whose parser validates enum and scalar prop values.

The exact generation and scoring condition is documented on the methodology page.

Data for agents and analysis