8 min read
Agent memory and RAG development: what to store, what to forget
Most agent memory problems are storage decisions made by default. Here is how to decide what an agent should remember, where each kind of memory belongs, and when retrieval is not needed at all.
Good RAG development starts with deciding what an agent should remember, not with choosing a vector database. Store durable facts that the user or your system owns, with their source and an expiry; retrieve them on demand; and forget raw conversation by default. Memory that is never pruned becomes stale context, a privacy liability, and an injection channel.
Retrieval-augmented generation, or RAG, means fetching relevant information at request time and giving it to the model as context. Agent memory is broader: it includes what the agent carries within a run, across a conversation, and across sessions. This guide separates those layers, explains where each belongs, and covers the cases where you need no retrieval system at all.
Why agent memory goes wrong
Memory feels like a single feature, "the agent remembers", so it often gets a single implementation: save everything and search it later. That blurs very different kinds of information. A customer's current plan belongs in your database, where it is always correct. A policy document belongs in a retrieval index, where it can be cited. A passing remark from last month's chat may not belong anywhere.
When those kinds are mixed, three problems appear. Retrieved context contradicts the system of record, because an old summary says one thing and the database says another. Personal data accumulates without a retention decision. And anything written into memory becomes something the agent will later read and trust, so a single manipulated input can influence behavior for months.
What teams often get wrong
- Storing every transcript. Full conversation logs are expensive to search, full of noise, and full of personal data. They rarely make better context than a few extracted facts.
- Embedding what should be queried. Order status, balances, and account settings live in structured systems. Retrieving them by semantic similarity from copied text invites stale or wrong answers.
- No provenance. A memory item without its source and date cannot be cited, verified, or corrected.
- Unreviewed memory writes. Letting the agent save whatever it concludes, from whatever it read, is how memory poisoning happens.
- Missing tenant filters. A shared index without strict per-customer filtering can surface one customer's documents to another.
- No deletion path. If a user asks to be forgotten, or a document is withdrawn, the index and memory store must forget too.
A practical approach: four layers of memory
Give each kind of information one home, with a clear write path and read path.
- Working context: what the agent needs during one run, such as the current request, recent tool results, and intermediate notes. It lives in the run's state and disappears when the run ends, apart from what you log for audit.
- Conversation history: recent turns in the current session. Keep a window of recent messages and summarize older ones when the session grows long. Checkpoints that let a run resume are part of this layer, not long-term memory.
- Long-term user memory: durable preferences and facts that the user has provided or confirmed, such as their role, their preferred format, or saved references. Store it as small, structured items with a source, an owner, and an expiry.
- Knowledge retrieval: documents the agent can search and cite, such as policies, product documentation, or contracts. This is classic RAG: an index built from source documents, re-indexed when the sources change.
Next to these sits your system of record. When the agent needs a fact that a database already holds, it should query that database through a tool, not remember a copy.
Decide per item
For anything you consider storing, answer five questions: who owns it, how sensitive it is, how long it stays true, who or what may write it, and how it will be retrieved. If you cannot answer them, do not store it yet.
An illustrative placement
Take a hypothetical account-support agent for a software product. The customer's subscription tier and invoice history stay in the billing system and are read through a tool on each request. The product documentation and refund policy go into a retrieval index, re-indexed when either changes, and every answer cites the section it used. The fact that this customer prefers short answers and works in the finance team becomes a long-term memory item, saved because they said so, with a review date. The details of the invoice they asked about today live only in the conversation and the run's working context. Nothing from the chat is copied into the index.
Placed this way, each fact has one authoritative home, and the question of what to delete when the customer leaves has a clear answer.
Write deliberately
Prefer explicit saves, where the user chooses to keep something, or extraction steps that propose memory items for confirmation. Treat any automatic memory write like an action: validate it, attribute it, and make it reversible.
Retrieve with scope and citations
Filter by tenant and user before similarity search, not after. Combine keyword and semantic search for better recall on names and codes, rerank the candidates, and pass the model a small number of passages with their sources. Ask the model to cite, and check that citations point to passages it was actually given.
Implementation considerations
- Chunking: split documents along their structure, such as headings, sections, and clauses, rather than fixed character counts, and keep the heading path with each chunk.
- Prompt budget: cap how many memory items and passages enter the prompt, and clamp each one's length. More context is not always better; irrelevant passages distract the model.
- Framing: label retrieved text and memory as reference data, not instructions, and never let it change permissions or tool access.
- Freshness: attach dates, re-index on source changes, and expire items whose truth decays, such as a user's current project.
- Retrieval evaluation: build a set of questions with known supporting passages and measure whether retrieval finds them, separately from whether the final answer is good.
- Deletion: make it possible to delete a user's memory and remove a document from the index, and verify that both actually happen.
Trade-offs
Large context windows make it tempting to skip retrieval and send everything. For small, stable corpora that works well and removes a moving part. For large or frequently changing corpora it raises cost per request and lowers answer quality, because relevant passages get buried.
A dedicated vector database scales well and adds another system to secure and operate. A vector extension in the database you already run is often enough for moderate volumes and keeps tenant filtering in familiar territory.
Automatic memory extraction makes an agent feel attentive and raises privacy and poisoning risk. Explicit saves are less magical and far easier to govern.
Lessons from ImadDhin work
These are code-level observations from this portal and its concept agents, not client outcomes.
- Memory is bounded twice. The Agent workspace stores saved memory items per session, capped at 50 items, with each stored body truncated. When building the prompt, only the most recent 24 items are included, and each body is clamped further.
- Memory is framed as data. The memory block in the prompt tells the model that saved items are user-provided reference data, not instructions, and that embedded instructions must never change permissions, tools, or system rules.
- Ownership is checked on every access. Reading or deleting a memory item verifies that it belongs to both the current session and the current visitor.
- Retention is a separate decision. That module decides limits and ownership but not how long items live beyond the session; retention belongs in an explicit policy rather than being implied by storage code.
- Small corpora go in the prompt. The eTROC concept agent bundles a few kilobytes of concept-deck text and passes it as reference context, capped in length, before any managed retrieval corpus is consulted. For a corpus that small, an index would add cost without adding quality.
- Run state is not memory. The workflow runtime stores checkpoints so runs can resume, with a size budget that rejects oversized state. Those checkpoints serve recovery, not long-term recall.
- Records beat recollection. The public FoCoCo case study describes a coach that connects conversations, practice, rounds, and reflection across a phone app and web app; that kind of continuity depends on the user's own structured records rather than on the model remembering past chats.
Common mistakes to test for
- Ask as one customer about another customer's documents and confirm nothing leaks.
- Save a memory item containing instructions and confirm the agent does not follow them.
- Change a source document and confirm answers reflect the new version after re-indexing.
- Delete a user's memory and confirm it no longer appears in any prompt.
- Ask a question whose answer lives in the database and confirm the agent queries it instead of citing an old copy.
- Check that every citation points to a passage the model was actually given.
When a simpler solution is better
If your reference material fits comfortably in the prompt, put it there and skip the index. If the answers live in a database, give the agent a query tool instead of building RAG over exported text. If users do not return across sessions, you may not need long-term memory at all. Add retrieval infrastructure when the corpus is too large or changes too often for simpler options. For the choice between retrieval, fine-tuning, and rules in chatbots specifically, see RAG vs fine-tuning vs rules.
Store less, retrieve better
Give each kind of information one home, write memory deliberately, retrieve with scope and citations, and make forgetting a feature. If you are designing memory or retrieval for an agent or AI product, see AI agent development and AI product development, or discuss your data in a 30-minute call.
Frequently asked questions
What is the difference between agent memory and RAG?
RAG retrieves relevant documents at request time and gives them to the model as context. Agent memory is broader and includes working context within a run, conversation history, and durable user facts across sessions. RAG is one way to implement part of it.
Should an AI agent store full conversation transcripts?
Usually not as memory. Keep recent turns for the current session, summarize older ones, and extract small confirmed facts for long-term memory. Retain full transcripts only where there is a clear purpose and a retention policy.
Do I need a vector database for RAG?
Not always. If the corpus fits in the prompt, send it directly. For moderate volumes, a vector extension in your existing database is often enough. A dedicated vector database makes sense at larger scale or with demanding search requirements.
How do I prevent memory poisoning?
Treat memory writes like actions: prefer explicit saves or confirmed extractions, record the source, frame retrieved memory as data rather than instructions, and make items reviewable and deletable.
How do I keep one customer's data out of another customer's answers?
Apply tenant and user filters before similarity search, store ownership with every chunk and memory item, and test cross-tenant queries explicitly before launch.
Design memory your agent can be trusted with
Walk through your data sources and decide what the agent should remember.
Book a 30-minute callScoped retrieval, cited answers, bounded memory, and deletion paths.
See AI agent developmentRetrieval and memory designed as part of the product, not bolted on.
Explore AI product developmentKeep reading
AI chatbot development: RAG vs fine-tuning vs rules
Train it on our data can mean three very different things. How to choose between rules, retrieval and fine-tuning for each kind of question your chatbot has to answer.
AI agent evaluation before production: a practical evaluation harness
An agent that looked good in five demo conversations can still fail on the sixth real one. A small, repeatable evaluation harness turns quality from an impression into a report you can rerun on every change.
AI agent permissions, tool scopes and audit trails
An agent's blast radius is the sum of everything it is allowed to call. Here is how to scope permissions per tool, keep credentials narrow and short-lived, and record an audit trail that answers who asked, what ran, and what changed.