Reading Agent Memories: Retrieval by Relevance, Recency, and Subject
Reading the right memories back at the right time is half of a memory system. When an agent starts a task, the relevant memories must be found and placed into context before it acts.
This is a retrieval problem — the same shape as RAG — with three signals a document store often lacks: relevance, recency, and subject.
Retrieval is where memory becomes useful
Writing memories is half the system; reading the right ones back at the right time is the other half. A memory the agent never recalls might as well not exist. When an agent starts a task, the relevant memories must be found and placed into context before it acts — and getting that retrieval right is what separates an agent that feels like it knows you from one that doesn't.
Memory retrieval is RAG on the agent's past
The mechanism is familiar. Embed the current situation — the query, the task, the recent context — and find the memories whose vectors are closest, optionally combined with keyword and metadata filters. Everything the RAG discipline teaches about hybrid search, reranking, and retrieving the right number of items applies directly. Memory retrieval is RAG pointed inward.
Relevance, recency, and subject
Memory retrieval has signals a document store often lacks. Recency matters — a preference expressed yesterday usually outweighs one from a year ago. Subject matters — retrieve memories about this user, this project. And importance matters — a correction should outrank a passing remark. The best memory retrieval blends semantic relevance with these signals rather than ranking on similarity alone.
A year-old note should rarely outrank yesterday's correction.
Retrieve for the situation, not just the query
What's relevant depends on the whole situation, not just the literal question. When a user asks a support question, the relevant memories include not only ones matching the words but ones about who they are, what they were doing last time, and preferences that shape a good answer. Retrieving well means embedding the situation — recent context, active task, and subject — rather than the query alone.
Retrieve wide, then narrow
As with any retrieval feeding a model, more is not better. Flooding the context with every vaguely related memory wastes budget and buries the memory that mattered. Retrieve generously, then narrow to the few genuinely relevant memories — the retrieve-wide, rerank-narrow pattern from RAG applies here exactly. A handful of precisely chosen memories beats a pile of loosely related ones.
ctx = store.retrieve(
subject='user_1837', # privacy + relevance
query=current_situation, # not just the literal question
top_k=5, # a few, not a flood
) # ranked by relevance + recency + importance
The cold-start problem
A memory system is least useful on its very first interaction with someone, when there's nothing stored to retrieve. This is expected, but it shapes design: the agent should degrade gracefully to sensible default behavior when memory is empty, rather than acting broken. A memory system earns its value over time.