Context Engineering Is the New Prompt Engineering
Prompts are a fraction of what the model sees. The real lever is everything in the context window — and most teams never audit it.

Prompt engineering got famous because it is visible: you write a string, the model answers. But on any non-trivial system the prompt is maybe 20% of what the model actually reads. The other 80% — retrieved passages, tool outputs, prior turns, system scaffolding — is the context. That is where the real quality lives, and almost nobody audits it.
Treat the context window like memory, not text
Every token competes for attention. A retrieved document that repeats the same point three times does not reinforce the point — it dilutes the signal and pushes the actual instruction toward the middle of the window where it gets less weight.
- Order matters: instructions at the very top and very bottom are attended to more than the middle.
- Recency matters: the last few thousand tokens dominate the next generation.
- Repetition is not emphasis; it is noise.
A simple context budget
Before shipping a feature, write down where every token comes from and cap each source:
system_scaffold: 800 tokens (hard cap)
retrieved_docs: 4000 tokens (top-k, reranked, deduped)
tool_outputs: 3000 tokens (summarised, truncated)
history: 2000 tokens (rolling, summarised older turns)
total budget: ~10000 tokens
If any source blows its budget, truncate or summarise before it enters the window. This is boring, mechanical work, and it is where most accuracy gains actually come from.
The shift in mindset
Prompt engineering asks: what do I tell the model? Context engineering asks: what does the model see, in what order, and how much of each? Once you start auditing the second question, the first one gets a lot easier.
You are not writing a prompt. You are assembling a context. The prompt is just the part you typed.



