Skip to main content

The Art of the Stream: A Pragmatist's Guide to Real-Time LLM Responses

This article is the 1st in a series of 5 blog posts...

The Art of the Stream: A Pragmatist's Guide to Real-Time LLM Responses

Introduction: Beyond the Blinking Cursor

This article is the 1st in a series of 5 blog posts on AI Chatbot UI and LLM Responses.

The user experience of a modern AI chatbot is often defined by how it delivers its response. There’s a significant difference between an application that shows a blinking cursor for several seconds before delivering a block of text, and one that reveals its response progressively, token by token. This streaming of a Large Language Model’s (LLM) output creates a powerful perception of speed and responsiveness. The user feels they are in a real-time dialogue, rather than waiting for a slow query to complete.

While the front-end effect is straightforward, the back-end architecture that enables it involves nuanced decisions. The choice of communication protocol—primarily between WebSockets and Server-Sent Events (SSE)—has lasting implications for system complexity, infrastructure cost, and scalability. This article provides a decision-making framework for this engineering challenge, dissecting the trade-offs between the primary streaming protocols and concluding with a tutorial for building a streaming backend on Amazon Web Services (AWS).

The Contenders: WebSockets and Server-Sent Events (SSE)

WebSockets: Full-Duplex Communication

The WebSocket protocol provides a full-duplex communication channel over a single, long-lived TCP connection. After an initial HTTP “upgrade” handshake, the connection allows for low-latency, event-driven data exchange initiated by either the client or the server at any time. This bi-directional capability makes WebSockets an excellent solution for applications requiring true, simultaneous two-way interaction, such as collaborative document editors or multiplayer online games. The protocol uses its own URI schemes, ws:// for unencrypted connections and wss:// for secure connections.

However, for a standard LLM-powered chatbot, this power can be a form of over-engineering. A typical chatbot interaction follows a predictable pattern: the user sends a prompt (client-to-server), and the LLM streams a response (server-to-client). The LLM rarely initiates a message without a preceding user request. Using a full-duplex protocol for what is essentially a half-duplex interaction introduces complexity. Developers must manually implement logic for connection state management, periodic keep-alive pings to prevent timeouts, and reconnection logic, as this is not a built-in feature of the WebSocket API.

Server-Sent Events (SSE): The Elegance of Simplicity

In contrast, Server-Sent Events (SSE) offer a simpler alternative designed specifically for uni-directional, server-to-client data streaming. As part of the HTML5 standard, SSE operates over conventional HTTP, using a single connection to push updates from the server.

SSE’s most compelling features are its built-in mechanisms for fault tolerance. Browsers that support the EventSource API will automatically attempt to reconnect if the connection is lost. Furthermore, if the server includes an id field with its events, the browser will send a Last-Event-ID header upon reconnection, allowing the server to resume the stream where it left off. This built-in support for reconnection and stream resumption makes SSE a robust choice for applications like live news feeds, stock tickers, and streaming LLM responses.

The protocol’s reliance on standard HTTP is a significant advantage in modern cloud infrastructure. Because SSE uses HTTP, it integrates seamlessly with existing web infrastructure, including load balancers, reverse proxies, and firewalls, which are highly optimized for HTTP traffic. WebSocket connections, using a distinct protocol, often require special configuration on these components. This alignment with HTTP also makes SSE a natural fit for serverless architectures like AWS Lambda. Managing the persistent, stateful connections of WebSockets in a serverless environment requires additional services, whereas SSE’s simpler nature avoids this overhead.

The Showdown: A Comparative Framework

To make an informed decision, a direct comparison of the protocols within the context of LLM streaming is essential.

**Criterion ****WebSockets **Server-Sent Events (SSE) Verdict for LLM Chatbots Communication Model Bi-directional, full-duplex. Client and server can send data at any time. Uni-directional, half-duplex. Only the server sends data to the client. SSE. Most chatbots are request-response, making bi-directional communication unnecessary overhead. Underlying Protocol Custom WebSocket protocol (ws://, wss://) initiated via an HTTP upgrade. Standard HTTP/1.1 or HTTP/2. Works over http:// and https://. SSE. Leverages existing HTTP infrastructure, simplifying deployment and improving compatibility. Scalability & Infra Requires stateful connection management on the server, complicating horizontal scaling. Stateless by nature. Scales easily with standard HTTP load balancing. SSE. A stateless design is a significant advantage for scalability, especially in serverless environments. Error Handling & Reconnection Manual implementation is required on the client side to detect drops and re-establish connections. Automatic reconnection is built into the browser’s EventSource API. SSE. Built-in resilience reduces client-side code complexity. Data Format Supports both UTF-8 text and binary data formats. Limited to UTF-8 text messages only. Tie. LLM token streams are text, making SSE’s limitation irrelevant for this use case.

While the analysis favors SSE for typical chatbot applications, the best protocol depends on the product’s long-term roadmap. If the future vision for the agent includes capabilities like proactive notifications or complex multi-agent systems, WebSockets become a more compelling choice. The pragmatic approach is to choose SSE for speed and simplicity today but to be prepared to adopt WebSockets if the application’s interactivity requirements evolve.

Workshop: Building a Streaming API on AWS

This section details the implementation of a WebSocket-based streaming API. While SSE is often the recommended choice, the WebSocket architecture is more instructive, covering concepts like connection management that are crucial for building more advanced systems.

Architecture Overview

The architecture uses a suite of AWS services to create a scalable, serverless WebSocket API:

  • Amazon API Gateway (WebSocket API): Serves as the front door for client connections, managing the connections and routing messages to backend services.

  • AWS Lambda: Provides the serverless compute layer. Three functions will handle the core logic: connecting, disconnecting, and processing messages.

  • Amazon DynamoDB: A NoSQL database used to store the connectionId of each client. This is necessary because Lambda functions are stateless.

  • ApiGatewayManagementApi: An AWS service used by the Lambda function to post messages back to a specific client via its connectionId.

The data flow is as follows: A client establishes a WebSocket connection through API Gateway. The $connect route is triggered, invoking a Lambda that saves the connectionId to DynamoDB. When the client sends a message, a custom sendmessage route invokes another Lambda. This function processes the message, calls an LLM, and streams the response back to the client’s connectionId. When the client disconnects, the $disconnect route triggers a final Lambda to clean up the connectionId from DynamoDB.

Step-by-Step Implementation

The implementation involves configuring the AWS resources and deploying the Lambda function code.

  • Resource Setup (AWS Console/CLI):

  • Create a DynamoDB table with a primary key of connectionId (string).

  • Create three Lambda functions (e.g., in Python or Node.js) with IAM roles that grant them permission to access the DynamoDB table and the ApiGatewayManagementApi.

  • In API Gateway, create a new WebSocket API. Define three routes: the predefined $connect and $disconnect routes, and a custom route named sendmessage.

  • Configure an integration for each route, pointing it to the corresponding Lambda function.

  • Deploy the API to a stage to get a connectable WebSocket URL.

  • Backend Logic (Lambda Code):

  • connectHandler: Receives the connection event, extracts the connectionId, and writes it to the DynamoDB table.

  • disconnectHandler: Receives the disconnect event, extracts the connectionId, and deletes the corresponding item from the DynamoDB table.

  • sendMessageHandler: This is the core logic. It receives the message from the user, invokes the LLM service with streaming enabled, and iterates through the response chunks. For each chunk, it uses the ApiGatewayManagementApi client and the post_to_connection method to send the data back to the connectionId.

  • **Client-Side Logic (JavaScript): **

  • Use the native WebSocket object in JavaScript to connect to the deployed API Gateway endpoint.

  • Use event listeners (onopen, onmessage, onclose, onerror) to handle the connection lifecycle and display the incoming stream of tokens.

  • Use the socket.send() method to send a message to the backend.

Conclusion: Choose Your Stream Wisely

The decision of how to stream LLM responses is a foundational architectural choice. For the majority of user-facing chatbots with a simple, turn-based dialogue, Server-Sent Events (SSE) offer a superior path. SSE’s simplicity, reliance on standard HTTP, and built-in resilience make it easier to develop and cheaper to scale. However, as AI systems evolve into more complex and proactive partners, the full-duplex power of WebSockets provides the necessary foundation for richer experiences. By understanding the trade-offs, engineering teams can make a pragmatic choice for today’s needs while building an architecture that is prepared for the future.

Topics:

Discuss this architectural approach.

Tell us what you'd like to build: an app, an agent, or a team alongside yours.