All articlesPrompt Engineering

Structured Output Engineering: JSON, Not Prose

When you need reliable, parseable output, you need structured generation. How to enforce schemas and stop fighting the model over formatting.

Sri Raman21 August 20268 min read
Structured Output Engineering: JSON, Not Prose

Asking a model to 'return JSON' and hoping it does is not engineering. Models add prose before the JSON, miss fields, invent fields, or wrap the JSON in Markdown code fences. The fix is not better prompting — it is structured generation.

Structured generation vs. prompting for JSON

Structured generation constrains the model's output to a grammar at decode time. It cannot produce a token that violates the schema. This is different from asking the model to 'please output valid JSON' — that is a request the model can refuse. With constrained decoding, the model physically cannot emit invalid output.

When to use it

Use structured generation whenever the output is consumed by code: tool calls, data extraction, classification with fields, API response formatting. Use free-form generation for content meant for humans: articles, summaries, chat responses. Mixing the two — asking for prose with a JSON block embedded — is where most formatting failures live.

Handling the 'I don't know' case

Your schema must include a null or 'unknown' option for every field. If the model cannot find a value, it needs a valid place to put 'I don't know.' Without that, it will hallucinate a value to fill the required field. Optional fields with null defaults are not a nice-to-have; they are a hallucination defense.

Validate after generation

Even with structured generation, validate the output against your schema on the server. Constrained decoding prevents malformed JSON, but it does not prevent semantically wrong values — a date in the past where a future date is expected, a number where a string ID is needed. The schema is the floor, not the ceiling.

Share this article