Skip to main content

Breaking the Text Wall: Building Rich, Interactive UIs with MCP-UI

Architectural pattern that solves the state synchronization challenges of ad-hoc UI.

Breaking the Text Wall: Building Rich, Interactive UIs with MCP-UI

Introduction: The Expressive Limits of Text

A purely text-based conversational interface, for all its natural language prowess, has expressive limits. Consider the task of purchasing a configurable product, such as a t-shirt, on a Shopify store. A text-only interaction can quickly become a tedious and error-prone back-and-forth: 

*User: “I’d like a t-shirt.” *

Agent: “What color?” 

User: “What colors are available?” 

Agent: “Red, blue, and green.” 

*User: “Show me blue.” *

Agent: “What size?” 

*User: “Large.” *

*Agent: “Large is out of stock in blue.” *

This clunky exchange highlights a fundamental truth: for many real-world tasks involving complex data, visual selection, or dependent options, a wall of text is an inefficient and frustrating medium. To deliver truly effective experiences, AI agents must be able to augment the conversation with rich, interactive graphical components.

The central engineering challenge is how to embed these components without the agent losing control of the conversational state and application logic. This article introduces the Model Context Protocol for UI (MCP-UI), an open-source protocol that provides a clean and robust solution to this problem.

The Ad-Hoc Anti-Pattern: Why Just Embedding HTML Fails

Faced with the need for richer UI, a developer’s first instinct might be to have the agent simply return a blob of HTML, CSS, and JavaScript, which the client application then renders inside a sandboxed iframe. While seemingly straightforward, this ad-hoc approach contains a critical architectural flaw: the state synchronization problem.

The core failure of this pattern lies in the creation of two independent sources of truth. The AI agent maintains an internal “belief” about the state of the application—for example, the items currently in the user’s shopping cart. When the developer embeds a self-contained HTML component, such as a product card with its own “Add to Cart” button, that component also has its own internal state. When the user clicks the button inside the iframe, the component’s JavaScript might update its own appearance and make a direct API call to the backend to modify the cart.

At this moment, the application state is fractured. The cart has been updated on the server, but the AI agent, which operates in a separate context, is completely unaware of this action. Its internal state is now stale. If the user then continues the conversation by typing, “Great, now add a two-year warranty,” the agent’s logic will likely fail. It may not know which product to apply the warranty to, or it may believe the cart is still empty. The conversational context is broken because the UI component “went behind the agent’s back.”

This reveals that the fundamental challenge is not simply rendering UI; it is maintaining a single, authoritative source of truth for the application’s state. In an agentic system, that authority must be the agent itself.

The MCP-UI Architecture: A Protocol for Collaborative UI

MCP-UI is an architectural pattern designed to solve the state synchronization problem by enforcing the agent’s role as the sole controller of state. It extends the base Model Context Protocol (MCP) by specifying a standard way for agents to deliver UI components and for those components to communicate back to the agent. This approach treats the embedded UI components not as independent mini-apps, but as “dumb” or stateless views that are responsible only for rendering data and emitting user intents.

Delivering the UI: Resources and Rendering

The protocol defines how an agent can specify a UI component to be rendered by the client. This is done via an embedded_resources block within the agent’s standard JSON response. MCP-UI supports three primary delivery methods:

  • Inline HTML: The agent provides the full HTML, CSS, and JS content directly in the response. The client can then render this in a sandboxed iframe using the srcDoc attribute. 

  • Remote Resources: The agent provides a URI that points to the UI component. The client then loads this URI into a sandboxed iframe. This is a common approach, allowing for complex components to be hosted independently. 

  • Remote DOM: For tightly integrated systems where sandboxing is not required, the protocol allows for the delivery of a DOM structure that can be rendered directly by the client-side application. 

The agent’s response payload specifies this component in a structured JSON format. The core of this is the UIResource interface, which defines the component to be rendered.

Here is a sample JSON response from an agent embedding a UI component:

{ 
  "response_text": "Here's the t-shirt you asked about.", 
  "embedded_resources": 
}

The client application is then responsible for interpreting this payload and rendering the component, often by creating an iframe and setting its src to the provided URI.

The Intent System: How the UI Talks Back

The most critical aspect of MCP-UI is its intent-based messaging system, which solves the state synchronization problem. When a user interacts with an embedded component—for example, by clicking a button or selecting an option—the component does not modify the application state directly. Instead, it uses the browser’s postMessage API to send a message to the host application (the chat client). This message contains a standardized JSON object describing the user’s “intent.”

This postMessage call is simple, standardized JavaScript.

JavaScript 

// Code inside the embedded iframe component 
const button = document.getElementById('add-to-cart-btn'); 
button.addEventListener('click', () => { 
  window.parent.postMessage({ 
    intent: 'add_to_cart', 
    payload: { 
      productId: 'B-123', 
      quantity: 1 
    } 
  }, '*'); // Target the parent window 
});

The protocol defines several standard intents, such as :

  • view_details: The user wants more information about an item. 

  • checkout: The user is ready to complete their purchase. 

  • notify: The component has performed a purely local action (like an animation) that the agent might need to be aware of. 

  • ui-size-change: The component’s dimensions have changed, and it needs the iframe to be resized. 

This architecture creates a closed loop of data flow that preserves the agent as the single source of truth:

  • The agent sends a response containing an MCP-UI component definition. 

  • The client application renders the component in an iframe. 

  • The user interacts with the component (e.g., clicks “Add to Cart”). 

  • The component’s JavaScript fires an intent message to the client, e.g., postMessage({ intent: ‘add_to_cart’, payload: { productId: ‘abc-123’, quantity: 1 } }, ’*’). 

  • The client’s JavaScript listens for these messages, captures the intent, and sends it to the agent as a new conversational turn. 

  • The agent receives the add_to_cart intent, processes it, updates its internal state (its “belief” about the cart’s contents), and formulates a response. This response might be a simple text confirmation (“I’ve added the t-shirt to your cart.”) or a new, updated MCP-UI component showing the current state of the cart. 

Workshop: MCP-UI in Action

Example 1: An Interactive Flight Booker

To illustrate this flow, consider a practical example of an interactive flight booking agent.

  • Step 1: The Agent’s Opening Move. A user types, “Find me flights from JFK to SFO for next Tuesday.” The agent processes this and responds with a JSON payload containing an MCP-UI remote resource pointing to a flight search component. 

  • Step 2: The Interactive Component. The client renders the component, which displays a form with fields for departure/arrival airports (pre-filled), date (pre-filled), and number of passengers. The component’s JavaScript is wired so that clicking the “Find Flights” button will gather the form data and post an intent message: postMessage({ intent: ‘search_flights’, payload: { from: ‘JFK’, to: ‘SFO’, date: ’…’, passengers: 1 } }, ’*’). 

  • Step 3: The User’s Intent. The user confirms the details and clicks the button. The client application catches the search_flights intent and sends it to the agent as a new message. 

  • Step 4: The Agent’s Rich Response. The agent receives the intent, calls a flight search API with the provided payload, and gets back a list of available flights. Its next response is not text, but another MCP-UI component: a visually rich carousel of flight options. Each card in the carousel displays the airline, times, and price, and includes a “Select this flight” button. Clicking this button would fire yet another intent, select_flight, continuing the collaborative workflow. 

Example 2: Embedding a Media Player

Let’s use the request for embedding a Spotify or SoundCloud player. This is a perfect example of a “remote resource” where the agent simply provides a URL to an existing web player.

  • Step 1: The Agent’s Response. The user says, “Play ‘Bohemian Rhapsody’ by Queen.” The agent finds the song and responds with an MCP-UI resource. 
{ 
  "response_text": "Sure, playing that for you now.", 
  "embedded_resources": 
}
  • Step 2: The Client Renders the iframe. The chat application receives this payload, sees the embedded_resources, and dynamically creates an iframe to render the component. The client-side JavaScript to handle this would be straightforward: 
JavaScript

// Code in the main chat application 
function handleAgentResponse(response) { 
  if (response.embedded_resources) { 
    const resource = response.embedded_resources.resource; 
    if (resource.mimeType === 'text/uri-list') { 
      const iframe = document.createElement('iframe'); 
      iframe.src = resource.uri; 
      iframe.sandbox = 'allow-scripts allow-same-origin'; // Set security sandbox 
      iframe.width = '300'; 
      iframe.height = '380'; 
      iframe.frameBorder = '0'; 
      document.getElementById('chat-messages').appendChild(iframe); 
    } 
  } 
}
  • Step 3: The Component Sends an Intent. While a standard Spotify embed might not send intents, a custom media player wrapper could. For example, a custom component could send a notify intent when the song finishes, allowing the agent to ask, “What would you like to play next?” This closes the loop and keeps the agent in control of the state. 

The Broader Implications: An “App Store” for Agent Capabilities

The conclusion of this technical deep dive is a forward-looking one. The standardization offered by protocols like MCP-UI has implications that extend far beyond simply making chatbots more interactive. History has shown that well-defined technical standards—from USB and HTTP to the ERC-20 token standard—are the catalysts for creating vast, interoperable ecosystems.

MCP-UI provides a standard interface for how AI agents can render and interact with graphical UI components. This standardization means that any agent that “speaks” MCP-UI can, in theory, use any component built to that same specification, regardless of who developed either piece of software. This paves the way for a future marketplace or “App Store” for agent capabilities, a concept that has already begun to emerge in agent-building communities. In this future, developers building a new travel agent won’t need to create their own flight selection UI from scratch. Instead, they can pull a best-in-class, pre-built MCP-UI flight selection component from a third-party marketplace, similar to how web developers today use component libraries like KendoReact or Ignite UI to accelerate their work. This will dramatically accelerate the development of sophisticated agents, allowing companies to focus on their core domain logic and reasoning capabilities while leveraging a rich ecosystem of standardized, interoperable UI components.

Topics:

Discuss this architectural approach.

Tell us what you'd like to build: an app, an agent, or a team alongside yours.