Designing Streaming LLM Responses That Feel Instant
Streaming is not just about speed — it is about perceived performance. How to chunk, display, and handle partial tool calls in a streaming UI.

Streaming responses make an LLM feel fast even when the total generation time is unchanged. The user sees the first token in 300ms instead of waiting eight seconds for a complete response. But streaming introduces a class of UI and engineering problems that batch responses do not have.
The partial-tool-call problem
When a model streams a tool call as JSON, the first chunks are incomplete JSON. You cannot parse '{"name":"sea' until the string closes. Two strategies: buffer until the JSON is parseable, or use a structured streaming parser that emits deltas. Buffering is simpler but adds latency. Delta parsing is faster but harder to get right and must handle truncated streams gracefully.
Markdown rendering during streaming
If you render Markdown as it streams, a half-written code block or an unclosed bold tag will produce a broken layout that flickers as more tokens arrive. The pragmatic fix: render plain text while streaming, then re-render as Markdown once the stream completes. The advanced fix: use a streaming-aware Markdown parser that handles unclosed delimiters.
- Show a typing indicator before the first token arrives.
- Auto-scroll only if the user is already at the bottom — do not yank them down if they scrolled up.
- Provide a 'stop' button that cancels the stream mid-generation.
Error handling mid-stream
If the stream fails halfway, the user has partial text. Do not discard it — show what was received, mark it as incomplete, and offer a retry. Silently truncating a response is worse than showing an error.




