Generative UI benchmark · Frameworks

Generative UI framework benchmark

How do OpenUI, Google A2UI, and Vercel json-render compare under the same briefs and generation condition?

What this benchmark answers

This is a controlled comparison of three generative UI formats: OpenUI, Google A2UI, and Vercel json-render. It covers 6 models, 46 briefs, and four generations per brief: 1,104 runs per format and 3,312 scored runs in total.

Structural validity and render success are separate. A generation can put something on screen while still failing the complete structural checks.

Format summary

Summary results for each generative UI format
FormatStructural validityRender successBlank screensRuns
OpenUI96.9%99.9%11,104
A2UI95.6%96.6%371,104
json-render82.8%99.6%41,104

Results by model and format

Structural validity, render success, and cost by model and format
ModelFormatValid runsValidityRender successCost / 46 screens
GPT-5.6 SolOpenUI183/18499.5%100.0%$3.08
GPT-5.6 SolA2UI177/18496.2%97.3%$6.84
GPT-5.6 Soljson-render152/18482.6%100.0%$6.30
Claude Opus 4.8OpenUI182/18498.9%100.0%$2.27
Claude Opus 4.8A2UI183/18499.5%100.0%$5.85
Claude Opus 4.8json-render161/18487.5%100.0%$4.88
Kimi K3OpenUI177/18496.2%100.0%$1.53
Kimi K3A2UI176/18495.7%97.3%$3.16
Kimi K3json-render131/18471.2%100.0%$2.91
Gemini 3.7 FlashOpenUI182/18498.9%100.0%$0.42
Gemini 3.7 FlashA2UI173/18494.0%96.7%$0.95
Gemini 3.7 Flashjson-render171/18492.9%100.0%$0.85
Qwen3.8 2.4TOpenUI169/18491.8%99.5%$0.81
Qwen3.8 2.4TA2UI168/18491.3%91.3%$1.90
Qwen3.8 2.4Tjson-render149/18481.0%97.8%$1.46
Muse Spark 1.2OpenUI177/18496.2%100.0%$0.71
Muse Spark 1.2A2UI178/18496.7%97.3%$1.38
Muse Spark 1.2json-render150/18481.5%100.0%$1.34

Controlled condition

  • One shared catalog of 70 equivalent components.
  • Each format uses the prompt generated by its own SDK.
  • The same two worked examples are carried by all three prompts.
  • Four generations per brief and a 16,384-token output ceiling.
  • Temperature 0.7 where supported, with minimal or no reasoning.
  • Each SDK’s shipped validation plus one identical shared completeness layer and component-count floor.

Every row is scored by one build of the shipped OpenUI parser, lang-core 0.2.16. See the complete methodology and reproduction steps.

Data for agents and analysis