Agent Memory Systems: Beyond Stuffing Everything Into Context
Long-lived agents need memory that survives across sessions. Working memory, episodic memory, and semantic memory — and when each fails.

An agent that forgets everything between runs is a stateless function call. An agent that remembers everything and stuffs it into the context window is a slow, expensive mess. Real agent memory is a system, not a variable — and it has layers.
The three layers
Working memory is the current context window — the active task, recent tool calls, the last few turns. It is fast, small, and volatile. Episodic memory is the log of past runs — what the agent tried, what worked, what failed. It is retrieved selectively when a similar task arises. Semantic memory is the distilled knowledge — facts, preferences, rules learned across all runs. It is the smallest and most valuable layer.
Episodic memory is RAG over your own history
When a new task arrives, the agent retrieves past runs that are similar. 'Last time I planned a deployment, I checked disk space first and that caught the issue.' This is RAG, but the corpus is the agent's own execution log, not external documents. The retrieval query is the current task description; the chunks are past run summaries.
Semantic memory is the hard part
Semantic memory — 'the user prefers concise responses,' 'API X is flaky on weekends' — must be written, not just retrieved. After each run, the agent decides: did I learn something that should persist? If yes, it writes to semantic memory. This is a write decision, and it is where memory systems go wrong: writing too much creates noise, writing too little loses the lesson.
The failure mode: memory drift
If semantic memory is never reviewed, outdated facts accumulate. 'The API returns XML' was true in January; in August it returns JSON. The agent keeps applying the stale rule. Memory needs a staleness check — a TTL, or a confidence that decays — so old memories do not override new reality.


