What is Intelligent UI?
Intelligent UI is the capability, introduced with GPT-6, that allows ChatGPT to respond with interactive interfaces rendered inline in the conversation (OpenAI announcement). A single response can combine prose with components such as sliders, forms, tables, charts, maps, and product cards. These components respond to input without a further model call, and they are drawn with ChatGPT's own design system rather than embedded as a web page.
The observations in this article were made on ChatGPT for the web in October 2026, using GPT-6 and GPT-6 Thinking.
Building blocks
ChatGPT's implementation divides the work between the model, OpenAI's server, and the client:
- Inference format: the model writes the interface as Markdown combined with JSX-like tags and JavaScript.
- Server-side compilation: the server converts each partial response into a JavaScript program and a JSON document of text and data.
- Client runtime and rendering: a sandboxed runtime executes the program, and ChatGPT draws the result with its native components.
- Design system and catalog: the components, properties, and design tokens available to the model.
1. Inference format
This is what the model writes. In ChatGPT, it is a language OpenAI calls DIL: Markdown for prose, JSX-like tags for components, and JavaScript for state and logic. We will follow one small response through every layer:
## Team plan estimate
Drag the slider to see the **monthly price** for your team.
{@body const [seats,setSeats] = DIL.useState(8)}
{@body const price = seats*29}
<box border padding={3} gap={2}>
<slider min={1} max={50} value={seats} onChange={setSeats}/>
<title size="xl">${price}/mo</title>
</box>The heading and the paragraph are ordinary Markdown. The tags are components from ChatGPT's catalog. The two {@body …} lines are JavaScript: the first declares a piece of state, seats, and the second derives price from it. The slider is bound to seats, so moving it updates the price. The full vocabulary is:
- Markdown
- component tags
{@body …}statements{expression}interpolation{#if}/{:else if}/{:else}conditionals{#each list as item, i}loops- event handlers
GenUIactions such asissueNewTurnandcopy
A dedicated format is needed because the model writes the interface token by token.
- It has to be easy to write reliably, so it is built from notation the model already knows well.
- It has to stay usable while half-written. Statements sit on their own lines, and any open element can be closed automatically. That lets the server cut a partial response at its last complete construct and still compile it.
1.1 Server-side compilation: script + data JSON
The client never executes the model's output as written. OpenAI's server compiles it into a JavaScript program and a JSON document, which are stored with the message (as model_dil_v2). The response compiles to this (formatted for readability):
function __dilSafe(evaluate, failureValue) {
try { return evaluate(); } catch { return failureValue; }
}
DIL.render(__dil.jsx(() => {
const __dilConstants = DIL.useConstants();
const __dilModelDataBindings = DIL.useAppData((appData) => appData.opGenui?.modelDataBindings ?? {});
const [seats, setSeats] = DIL.useState(8, { key: "seats" });
const price = __dilSafe(() => seats * 29, undefined);
return __dil.jsx(__dil.Fragment, null,
__dil.jsx("title", { size: "lg" }, __dilConstants["0"]),
__dil.jsx("text", null, __dilConstants["1"], __dil.jsx("bold", null, __dilConstants["2"]), __dilConstants["3"]),
__dil.jsx("box", { border: true, padding: 3, gap: 2 },
__dilSafe(() => __dil.jsx("slider", { min: 1, max: 50, value: seats, onChange: setSeats }), null),
__dil.jsx("title", { size: "xl" }, __dilConstants["4"], __dilSafe(() => price, null), __dilConstants["5"])));
}, { key: "body:2" }));{
"constants": {
"0": "Team plan estimate",
"1": "Drag the slider to see the ",
"2": "monthly price",
"3": " for your team.",
"4": "$",
"5": "/mo"
},
"appData": { "opGenui": { "componentResults": {}, "modelDataBindings": {} } }
}The Markdown is compiled into the same tree as the components. The heading becomes a title, the paragraph a text with a bold inside it, and their words move into the constants table.
Compilation does work that every client would otherwise have to repeat:
- Plain function calls. Markup becomes calls to
__dil.jsx, so a JavaScript runtime can evaluate the program without a parser for DIL. - Error isolation. Expressions are wrapped in
__dilSafe, so an expression that throws removes one element instead of aborting the whole render. - Text in a separate table. Static text moves into the constants table, so as a response streams, growing text changes the data rather than the program.
- Stable state keys. Each piece of state receives a key (
{ key: "seats" }), so its value survives every recompilation. - Repair and validation. Incomplete statements and tags are dropped, unclosed elements are closed, and properties that fail validation against the catalog are removed and recorded as diagnostics.
The JSON document holds the text constants and any data the server resolves for the response, such as image search results (see Data).
2. Client side
The client receives the compiled program and the JSON document. Its work is split between a runtime, which executes the program, and a renderer, which draws the result.
2.1 JavaScript runtime
The program is model-written code, so it does not run in the ChatGPT page. ChatGPT loads a hidden iframe (runner.html), sandboxed with allow-scripts and a content security policy of default-src 'none', which starts a Web Worker.
- Lockdown. Before evaluating a program, the worker removes network access, timers, messaging, and dynamic code evaluation from its global scope, and freezes the remaining globals.
- Evaluation. It then evaluates the program with
new Function. The runtime objects (DIL,__dil,GenUI) and the catalog's composite components are passed in as parameters. - Watchdog. A program that does not respond within a timeout is quarantined, and the worker is restarted.
The runtime is a small reconciler in the style of React. It renders the component and keeps hook state in keyed slots. It then compares the resulting tree with the previous one and encodes the differences as a list of operations. It does not draw anything.
The following illustrative example shows the operations from a first render, with one line per node. Entries that list each element's property names are omitted:
CREATE #1 title SET size = "lg" PLACE under root at 0
CREATE #2 text "Team plan estimate" PLACE under #1 at 0
CREATE #3 text PLACE under root at 1
CREATE #4 text "Drag the slider to see the " PLACE under #3 at 0
CREATE #5 bold PLACE under #3 at 1
CREATE #6 text "monthly price" PLACE under #5 at 0
CREATE #7 text " for your team." PLACE under #3 at 2
CREATE #8 box SET border = true, padding = 3, gap = 2 PLACE under root at 2
CREATE #9 slider SET min = 1, max = 50, value = 8, onChange = fn#1 PLACE under #8 at 0
CREATE #10 title SET size = "xl" PLACE under #8 at 1
CREATE #11 text "$" PLACE under #10 at 0
CREATE #12 text "232" PLACE under #10 at 1
CREATE #13 text "/mo" PLACE under #10 at 2Functions never leave the worker; the slider's handler is sent only as an identifier (fn#1). On the wire, the operations are encoded as a binary sequence of integers, with strings held in a separate table.
2.2 Rendering
The ChatGPT page applies the operations to its own component tree. Each CREATE instantiates a native component from ChatGPT's design system, and the page animates changes as they arrive. The page accepts operations only for known component types, so model output cannot introduce arbitrary markup or styles (with some exceptions).
Interaction runs in the opposite direction. When the user drags the slider to 9, the page sends the handler's identifier and arguments to the worker. The worker calls setSeats(9), re-renders, and returns update operations. The changes to the slider and price look like this:
page → worker: trigger callback_1 [9]
worker → page: SET #9.value = 9
TEXT #12 = "261"No model call is involved.
The operation protocol does not depend on the platform. The worker's sandbox also permits the globals of Hermes, the JavaScript engine used by React Native. This suggests that ChatGPT's mobile apps run the same runtime and apply the operations with their own native renderers; we have not verified this directly.
3. Design system and catalog
The catalog defines what the model can request. It is needed because the model does not build an interface from raw layout and styling rules. It chooses from components ChatGPT already knows how to draw, and styles them with design tokens such as padding={3}. Raw CSS values, such as pixel widths and hex colours, are accepted for some properties, but design tokens are preferred. As a result:
- Generated interfaces look like the rest of ChatGPT on every platform.
- The compiler has a schema to check output against. A property that does not exist on a component, or a literal of the wrong type, is removed during compilation and recorded as a diagnostic.
In one response we captured, the compiler removed two properties: fill on an icon (an unknown_prop diagnostic) and gap="1" on a box (an invalid_literal diagnostic).
The catalog has three parts:
- Native components. Around 70 components are defined in the component registry in ChatGPT's client code; 39 of them appear in the responses we captured.
- Design tokens for spacing, radius, colour, and size.
- Composite components written by OpenAI in DIL and sent to the sandbox prebuilt, such as the image and product components. In our captures, the model used these components but never defined its own.
Every part of the response maps to a catalog entry:
| In the response | Catalog entry | Resolved value |
|---|---|---|
## Team plan estimate | title | lg size token |
**monthly price** | bold inside text | Inline emphasis |
box border padding={3} gap={2} | Layout container | Border, 12 px padding, 8 px gap (4 px spacing scale) |
slider min max value onChange | Input | ChatGPT's slider |
title size="xl" | Heading text | xl size token |
Streaming
Streaming text is simple: each new token is appended to what is already on screen. Streaming an interface is harder, for three reasons:
- The output is usually not runnable yet. At most moments it is an incomplete program, with a tag or expression still open, and it cannot be executed as written.
- The interface has to keep working while it grows. Components the user has already touched must keep their state.
- Some content arrives separately. Data such as images comes from the server, not from the text.
A plain stream of appended tokens cannot express this. ChatGPT instead streams patches to a structured message that holds the raw text, the compiled program, and its data side by side.
The response reaches the browser over a server-sent event stream (POST /backend-api/f/conversation). Each event is a JSON-Patch-style update to the message being built. A single event usually updates the raw DIL text and its compiled form together. This is one update from a captured response, shortened:
{"o": "patch", "v": [
{"p": "/message/content/parts/0", "o": "append", "v": " Sunday lamb roast with friends — generous food, …"},
{"p": "/message/metadata/model_dil_v2/code", "o": "replace", "v": "DIL.render(__dil.jsx(()=>{…"},
{"p": "/message/metadata/model_dil_v2/constants", "o": "append", "v": {"0": "Here's a plan for a proper Sunday lamb roast with friends — …"}},
{"p": "/message/metadata/model_dil_v2/constants", "o": "append", "v": {"1": "Since you're"}},
{"p": "/message/metadata/model_dil_v2/fallbackMarkdown", "o": "append", "v": " Sunday lamb roast with friends — …"}
]}The server does not compile incrementally. Every few hundred milliseconds, most likely with each new chunk of model output, it recompiles everything the model has written so far and sends the result. Compilation starts with the first token, before any tag has appeared.
Compiling a half-written response
At any moment, the model may be in the middle of a tag or an expression. The compiler cuts the source at its last complete construct:
- an unfinished
{@body}statement or tag is dropped; - an open element that already has content is closed automatically, and one without content is dropped;
- partial text is kept as it is.
The compiler records each cut in recoveryDiagnostics (unterminated_tag, unclosed_block, unterminated_braced_value). The entry disappears once the source is complete again. In the interface, text streams word by word, while each component appears only once its tag is complete.
Three kinds of update
We captured one response of 6,341 characters, delivered in 84 updates over 19 seconds. Apart from the updates that created the message, delivered image results, and marked completion, they fell into three groups:
| Update | Count | What is sent |
|---|---|---|
| Structure changed | 52 | Appended text and the entire compiled program |
| Only text grew | 14 | Appended text and an updated constant; the program is unchanged |
| Inside an unfinished tag | 10 | Appended text only; the interface does not change |
The compiled program is a single nested expression whose closing brackets change with every recompile, so it cannot be appended to and is replaced in full. Resent code made up 83% of the roughly 275 KB of patches for this response. Updates arrived at a median interval of 230 ms. In another capture, the interval was about 410 ms.
On the client
The page passes each new program to the sandboxed worker. The worker evaluates it, re-renders with the existing state, and sends update operations to the page. State keeps its values across recompiles because of the keys added during compilation (Section 1.1). If a new program fails to evaluate or render, the worker keeps the last one that worked.
The page then animates each change:
- text fades in over 0.7 s;
- new rows and grid items slide in over 0.42 s;
- charts draw over 1.8 s;
- container heights transition instead of jumping.
Data
Tool results contain values a user may act on: a retailer's price, a place's address, or an image from a page. Passing those values through the model risks copying errors or invented details. In our captures, ChatGPT supplied some data separately from the model's output. We observed field binding for product results, while weather figures from a web search were written directly into the response by the model. Where data was supplied separately, the model wrote a request or a reference, and the server supplied the data with the compiled program in appData. We observed two mechanisms.
Server-defined components
Some components are resolved by the server. To show an image, the model describes it instead of linking to it:
<AsyncImage query="slow roasted rosemary garlic lamb shoulder browned golden roast in baking tray" aspectRatio="5:4" maxWidth="152px"/>Once the tag is complete, the server runs an image search, checks the resulting URLs, and patches the result into the response, typically a second or two later:
"a588d0a9-…": {
"componentName": "AsyncImage",
"state": { "is_loading": false, "images": [{ "content_url": "https://images.openai.com/…", "url": "https://anecraefficientia.pt/…" }] }
}In the image-search responses we captured, image URLs were supplied by the server rather than written by the model. Some other components are resolved the same way:
AsyncImageGroup, for image carousels;Entity, for product and place chips;Cite, for source links.
Binding tool results
Results from the model's tool calls, such as a web search, are given IDs. The model refers to their fields by ID instead of retyping the values:
<title size="lg">{turn448791product0.displayed_price}</title>
<caption>{turn448791product0.merchant}</caption>The server supplies the referenced fields with the response:
"turn448791product0": {
"title": "Sony WH-1000XM6 Wireless Noise-Canceling Headphones",
"displayed_price": "₹35,989",
…
"merchant": "Amazon.in",
…
}The price on screen comes from the search result itself, not from the model's copy of it.
Actions (buttons, forms, and more)
Most interactions never leave the response. Dragging a slider or ticking a checkbox changes state inside the worker, and the page receives only the resulting operations (Section 2.2). Actions are the interactions that reach outside the response. The program reaches them through a small GenUI object provided by the host, with functions like:
issueNewTurn(text)copy(text)openUrl(url)openEntityDetail(ref)
Continuing the conversation
The only way an interface communicates with the model is issueNewTurn. It sends a new user message, and the program builds that message's text from its state:
<button block onClick={()=>GenUI.issueNewTurn("Develop a detailed dog vest design concept based on my moodboard. Style: "+style+". Main color: "+color+". Features: "+features.join(", ")+". Show visual design inspiration and practical construction ideas.")}>Develop this direction <icon name="arrow-right" inline/></button>After the user picked a style and a colour, the button produced this message:
Develop a detailed dog vest design concept based on my moodboard. Style: cozy. Main color: sand. Features: Reflective details, Adjustable straps. Show visual design inspiration and practical construction ideas.The message is stored exactly like a typed one, with nothing marking it as coming from the interface. Forms work the same way: <form onSubmit={…}> together with <button submit> calls issueNewTurn with the form's values.
Escape hatch: AppBlock, an app in an iframe
Some requests call for things the native components are not designed for, such as a drum machine that synthesizes sound with Web Audio. For these, the model can write an AppBlock: a self-contained web app in HTML, CSS, and JavaScript, embedded in the response. This is the start of one, shortened:
<AppBlock title="Drum Lab" icon="app-chatgpt" variant="inline" app_block_id="drum-lab-01">
<div id="dl" class="w-full min-w-0 space-y-4 text-base">
<style>
#dl{color:var(--viz-text)}#dl button{touch-action:manipulation}#dl .panel{background:var(--viz-panel);border:1px solid var(--viz-border);border-radius:15px}…
</style>
…
<button id="dl-play" class="btn" style="background:var(--viz-text);color:var(--viz-card);min-width:100px">▶ Play</button>
…
</div>
<script>
(function(){
const root=document.getElementById('dl');if(root.dataset.init)return;root.dataset.init="yes";
…
function audioInit(){if(!audio){const C=window.AudioContext||window.webkitAudioContext; if(!C)return false;audio=new C();…
…
})();
</script>
</AppBlock>An AppBlock is handled differently from the rest of the response. The compiler does not translate it. The compiled program contains only a placeholder (__chatgptClientDefinedWidget), and the client reads the app's source from the raw response text.
ChatGPT renders the app in a visible iframe on a separate domain. It uses the same host ("Skybridge") that it uses for apps, and writes the model's HTML into an inner frame:
<iframe sandbox="allow-scripts allow-same-origin" referrerpolicy="no-referrer"
src="https://codex-inline-visualization-<hash>.web-sandbox.oaiusercontent.com/?app=skybridge&deviceType=desktop&…">
<iframe id="root" src="about:blank"
sandbox="allow-scripts allow-same-origin allow-popups allow-popups-to-escape-sandbox allow-forms">
<!-- the app's HTML and script are written here -->
</iframe>
</iframe>An app in an iframe sits outside ChatGPT's design system. Left alone, it would look like a foreign web page in the middle of the conversation. The host prevents this in three ways:
- Shared base styles. It injects Tailwind's base styles, so the app starts from the same typography and spacing defaults as the page around it.
- Shared colours. It provides theme variables (
--viz-text,--viz-panel,--viz-accentand others) that the model's CSS refers to. The app uses these variables instead of hard-coded colours, so it follows ChatGPT's light and dark themes. - Sizing. It sizes the iframe to the app's content, so the app reads as part of the response rather than a scrolling box inside it.
Unlike a native response, the app draws its own interface. Its state lives in its own JavaScript variables, and in our captures it made no calls back to ChatGPT.
Markdown fallback
A response is stored once, but it can be opened from clients that cannot run it, most likely older app versions. For these, the server generates a plain-Markdown version, fallbackMarkdown, alongside every compiled program. It keeps the prose. Images, products, citations, and maps are written in the inline markup that ChatGPT clients already render. Interactive parts are reduced to static text.
Rough edges
- State resets on reload. Values the user changes are lost when the page is reloaded, even though the client sends them to a view-state endpoint.
- The model does not see interface state. It receives only the text of messages sent with
issueNewTurn. When asked, it could not report values the user had changed. - Follow-ups regenerate the interface. Each follow-up produces a new response with a new program; the existing interface is never modified.
- Most of the stream is resent code. In one response, resent programs made up 83% of 275 KB of patches for 6.3 KB of text.
- Interactions resend unchanged properties. A single click in a pricing widget produced 158 operations, only 4 of which changed anything.
- Logic errors are silent. An expression that throws an error renders nothing, and incorrect logic is not detected. In one dashboard, switching a chart's metric changed a revenue figure from $6,930 to $256,410.
- No apparent feedback on invalid properties. The compiler removes invalid properties and records diagnostics, but the model repeated the same mistakes in later responses. This suggests that those diagnostics may not reach the model.
- The fallback omits computed content. Values derived from state and content inside
{#each}do not appear in the Markdown version. - Image searches can miss. The model writes a query but never sees the result. A query for a close-up of a handlebar shifter returned an image from an electric bicycle listing.
- Layouts are fixed. Responses use fixed column counts and pixel widths, and none of the captured responses used the available breakpoint hooks.
AppBlockapps are isolated. Their state lives in their own JavaScript variables and is lost on reload, and they made no calls back to ChatGPT.
Methodology
All observations come from our own ChatGPT accounts, from the traffic the ChatGPT web app generates, and from the JavaScript that chatgpt.com serves publicly. They were made in October 2026 with GPT-6 and GPT-6 Thinking.
- Conversation exports. We exported conversations as the JSON that the web app loads for a conversation. For each response with an interface, the export holds the model's raw DIL, the compiled program, the constants and data, the Markdown fallback, and any compiler diagnostics.
- Stream captures. We recorded the server-sent event stream for several responses by wrapping the page's stream reader in the browser. We then replayed the patches to rebuild the message after every update.
- Browser inspection. We used the browser's developer tools to examine the page structure, including the iframes used for
AppBlockand the scripts the worker evaluates. - Client code. We read the sandbox runner and worker (protocol version 14), and the chatgpt.com bundles that hold the component registry, the renderer, and the widget host. We then ran the unmodified runner in a local test page to record the exact operations it emits.
- Compiler reimplementation. We reimplemented the server compiler from pairs of DIL source and compiled output. Its output matches OpenAI's byte for byte, both on the intermediate versions of a streamed response and on finished responses.
- Probes. We wrote prompts designed to exercise specific features: forms, lists that can be edited, charts, maps, timers, and buttons that send messages back.
Mobile implementations were not examined.