All articlesLLM Ops

Designing Streaming LLM Responses That Feel Instant

Streaming is not just about speed — it is about perceived performance. How to chunk, display, and handle partial tool calls in a streaming UI.

Sri Raman28 August 20268 min read
Designing Streaming LLM Responses That Feel Instant

Streaming responses make an LLM feel fast even when the total generation time is unchanged. The user sees the first token in 300ms instead of waiting eight seconds for a complete response. But streaming introduces a class of UI and engineering problems that batch responses do not have.

The partial-tool-call problem

When a model streams a tool call as JSON, the first chunks are incomplete JSON. You cannot parse '{"name":"sea' until the string closes. Two strategies: buffer until the JSON is parseable, or use a structured streaming parser that emits deltas. Buffering is simpler but adds latency. Delta parsing is faster but harder to get right and must handle truncated streams gracefully.

Markdown rendering during streaming

If you render Markdown as it streams, a half-written code block or an unclosed bold tag will produce a broken layout that flickers as more tokens arrive. The pragmatic fix: render plain text while streaming, then re-render as Markdown once the stream completes. The advanced fix: use a streaming-aware Markdown parser that handles unclosed delimiters.

  • Show a typing indicator before the first token arrives.
  • Auto-scroll only if the user is already at the bottom — do not yank them down if they scrolled up.
  • Provide a 'stop' button that cancels the stream mid-generation.

Error handling mid-stream

If the stream fails halfway, the user has partial text. Do not discard it — show what was received, mark it as incomplete, and offer a retry. Silently truncating a response is worse than showing an error.

Share this article