Structured Output Engineering: JSON, Not Prose
When you need reliable, parseable output, you need structured generation. How to enforce schemas and stop fighting the model over formatting.

Asking a model to 'return JSON' and hoping it does is not engineering. Models add prose before the JSON, miss fields, invent fields, or wrap the JSON in Markdown code fences. The fix is not better prompting — it is structured generation.
Structured generation vs. prompting for JSON
Structured generation constrains the model's output to a grammar at decode time. It cannot produce a token that violates the schema. This is different from asking the model to 'please output valid JSON' — that is a request the model can refuse. With constrained decoding, the model physically cannot emit invalid output.
When to use it
Use structured generation whenever the output is consumed by code: tool calls, data extraction, classification with fields, API response formatting. Use free-form generation for content meant for humans: articles, summaries, chat responses. Mixing the two — asking for prose with a JSON block embedded — is where most formatting failures live.
Handling the 'I don't know' case
Your schema must include a null or 'unknown' option for every field. If the model cannot find a value, it needs a valid place to put 'I don't know.' Without that, it will hallucinate a value to fill the required field. Optional fields with null defaults are not a nice-to-have; they are a hallucination defense.
Validate after generation
Even with structured generation, validate the output against your schema on the server. Constrained decoding prevents malformed JSON, but it does not prevent semantically wrong values — a date in the past where a future date is expected, a number where a string ID is needed. The schema is the floor, not the ceiling.



