Context Window Management: What to Keep, What to Evict
A full context window is not a feature — it is a liability. Strategies for prioritizing, compressing, and evicting context in long agent runs.

Context windows have grown from 4K to 128K to a million tokens. This has not made context management easier — it has made the failure mode worse. A full context window is slower, more expensive, and more prone to the 'lost in the middle' problem where the model ignores information in the center of the window.
Not all context is equal
System instructions are high-trust, stable, and must stay. The current user query is high-relevance and must stay. Retrieved documents are medium-relevance and may be dropped if they do not support the answer. Prior tool outputs are low-relevance once consumed — their result has been incorporated into the plan. Old conversation turns are the first candidates for eviction.
Summarize, do not truncate
When context overflows, the naive fix is to drop the oldest turns. This loses information. The better fix: summarize the oldest turns into a compact paragraph that preserves decisions and facts, then drop the originals. A 5000-token conversation summarized to 500 tokens retains the gist.
Tool output compression
A tool that returns 3000 tokens of JSON is rarely needed in full after the model reads it. After the model incorporates the result, replace the full output with a one-line summary: 'Retrieved 12 records; 3 matched filter X.' This keeps the window lean for future steps.
The priority rubric
Score each context chunk on recency, relevance to the current step, uniqueness, and authority. Keep the top-N by combined score; evict the rest. This is not a one-time decision — it runs at every step as the task evolves.



