Memory for Agents: Context, State, and Recall
Agents need to remember the right things at the right time. That is a design problem, not a database problem.
By NeuralNetworki.ng Team · AI Engineers
Memory is selective, not total
The intuitive goal is to give an agent perfect recall, to let it remember everything it has ever seen. It feels like more memory must mean more intelligence. In practice the opposite is often true. Stuffing everything into the context window makes the agent slower, more expensive, and, counterintuitively, less accurate, because the relevant facts get buried under a mountain of marginally relevant ones and the model's attention is spread thin.
Good agent memory is therefore not a storage problem; it is a retrieval and curation problem. The question is never "how do I remember all of this?" but "how do I surface the few things that matter for the decision in front of me, right now?" Get that framing right and the rest follows.
Three kinds of memory
It helps to separate memory into three layers, because each is built differently.
Working memory is the current task in flight: the goal, the last few steps, and the most recent tool results. It lives directly in the context window and should be kept deliberately lean. This is the agent's scratchpad, and a cluttered scratchpad slows everything down.
Short-term memory spans the current session, what the user said a few minutes ago, the shape of the conversation so far. Here the mistake is keeping the full raw transcript. A running summary that captures the gist, decisions, and open threads almost always serves the model better than a verbatim log, and costs a fraction of the tokens.
Long-term memory persists across sessions: a user's preferences, the outcome of past projects, durable facts about their domain. This cannot live in the context window, there is far too much of it, so it lives in external storage and is pulled in on demand. This is where retrieval comes in.
Retrieval, used carefully
The standard approach to long-term memory is to store facts as embeddings in a vector store and, at each step, retrieve the items most similar to the current situation. Done well, it gives the agent the feeling of a long memory without the cost of carrying it all.
The dominant failure mode is retrieving too much. Pull back twenty loosely related chunks and you have reintroduced the very noise you were trying to avoid, now the model has to find the signal among low-relevance passages, and it often does not. Tune for precision over recall: fewer, more relevant items beat a comprehensive dump almost every time. And always carry through provenance, where each retrieved fact came from, both so the model can weigh its reliability and so you can debug when it acts on the wrong one.
Summarise to survive
Any agent that runs for more than a handful of steps will eventually drown in its own history. The cure is active compression. Periodically fold older turns into a concise summary and discard the raw text behind it. The window stays manageable and the cost stays bounded.
The skill is in what you choose to keep. The valuable residue of past steps is not the blow-by-blow, it is the decisions that were made, the constraints that were discovered, the dead ends already ruled out, and the questions still open. Keep those. Drop the intermediate tool output that has already been digested. A good summary lets the agent continue as if it remembered everything important, because, functionally, it does.
Forgetting is a feature
There is a quiet assumption that more memory is always safer. It is not. Stale memory is often worse than no memory, because the agent acts on it with full confidence.
Any fact that can change over time, an address, a price, an account balance, a ticket status, should carry a freshness window. Past that window, the agent should re-fetch the current value rather than trust what it cached. An agent that confidently quotes last week's price, or acts on a status that has since changed, does more damage than one that simply looks it up again. Designing deliberate forgetting, knowing which facts are durable and which are perishable, is as important as designing what to remember. The best memory systems are not the ones that hold the most; they are the ones that consistently surface the right thing, fresh, at the right moment.
Related work
This is the kind of problem we solve in Agentic AI Systems. See it in practice in our Agentic Honeypot, ARGUS case study.
Talk to us about your project