Better Maps for an Agent’s Memory
Repeated work can train a durable access layer around an unchanged model: links, rankings, and reference artifacts that make later context easier to find. The useful unit is not memory volume but transferable retrieval experience, and that experience should be evaluated separately from the underlying record.
In one recent experiment, the researchers changed a document store and left the reader alone.
The reader was the same model, with the same tools and the same limit on how many actions it could take. What changed was the material around it. A separate training agent had answered questions, seen the correct answers, and then added links and index documents to the store. The researchers froze snapshots of that changing store and repeatedly sent the unchanged reader back in.
On questions the store had already trained on, the reader eventually used 31 percent fewer actions while answering more accurately. The model had not been fine-tuned. The reader’s instructions had not improved. Something else in the system had learned.
That sounds, at first, like an elaborate way of saying that memory helps. But the experiment became more interesting one step away from the training questions.
When a new question used the same template and both of its relevant keys had already been indexed, the action saving disappeared. Accuracy, however, transferred farther: the paper reports an F1 gain of 0.167 when both relevant keys had been covered during training, 0.100 when one had been covered, and no gain when neither had been covered.
The system had learned two different things with two different ranges. It had learned some reusable structure about the documents. It had also learned shortcuts that were useful mainly when the exact question came back.
That distinction suggests a more precise way to think about long-lived agent systems. They can remember what happened before. They can also remember what previous work taught them about where to look next.
Those are not the same asset.
The reader did not get smarter
The study, Training a Knowledge Base, was posted on August 22 and is listed as submitted to IEEE BigData 2026. Its main experiment uses a generated fictional corpus, and its second arm uses PhantomWiki, another synthetic world. The authors did not run a real-text arm, and the main training run used a single seed. So this is not yet evidence that a production knowledge base will improve in the same way.
What the setup does unusually well is isolate the moving part.
The trained store contains 1,913 task-conditioned links. An unsupervised entity index used for comparison contains 196,112. The authors report that the much smaller trained structure produces more accuracy and action saving per portion of the corpus it reaches. Because the reader is held fixed across the store snapshots, the experiment can attribute the difference to changes in external state rather than to a changing reader.
But the two endpoints behave differently. The largest efficiency gain appears only after two curation passes over the same training questions. The paper is explicit that indexing the right keys was not enough to preserve that saving on a different question. Accuracy generalized by coverage; efficiency mostly did not.
That is exactly the sort of result that gets blurred by the word memory. A store can contain the same underlying facts while gaining a new layer of organization over them. Some of that organization can generalize. Some of it behaves more like a cache.
The useful question is therefore not simply how much history an agent retains. It is what durable structure the system derives from using that history, and how far that structure travels before it stops helping.
A retriever can keep its receipts
A second experiment reaches the same distinction without editing the underlying memory at all.
Many memory systems first retrieve a broad set of possible memories and then ask a language model to judge which ones are actually relevant. That second step can be expensive. Worse, a stateless reranker throws the judgments away. A later question may force the system to pay again to learn that the same old memory is useful—or useless.
An August 24 preprint, The Retriever Should Remember, asks what happens if those relevance judgments become persistent state.
The researchers’ system keeps a sparse table of query-to-memory scores. As new questions arrive, it directly scores only some candidate memories and estimates the rest from the structure in earlier scores. Crucially, the experiment does not let the memory store itself grow. For each conversation in the LoCoMo benchmark, the complete memory store is built first and then held fixed. Only the record of retrieval judgments accumulates.
That makes the result unusually clean. The system is not getting more memories. It is getting more experienced at navigating the same memories.
With ten memories allowed into the final answer context, the best reported configuration raises the paper’s language-model-evaluated answer accuracy from 82.21 percent with ordinary semantic retrieval to 88.83 percent—an increase of 6.62 percentage points. Across the full evaluation, it makes 78,736 direct reranker calls, compared with 307,982 for reranking every candidate. In the final experience stage, when only 17.5 percent of each 200-memory candidate pool is directly scored, its top-10 accuracy remains 5.53 points above semantic retrieval.
The experiment has its own boundaries. It is one conversational-memory benchmark. The memory identities are stable. GPT-4o-mini generates the answers and, in a separate call, also evaluates them against the references. The study does not show what happens when memories themselves are constantly being added, deleted, or rewritten.
Still, it demonstrates something the first experiment only implied. An expensive decision made during retrieval can itself become reusable memory.
A system can preserve not only the old conversation, but also the fact that an earlier question taught it which parts of that conversation mattered.
The production version is messier
The same idea appears in a much less controlled setting: a large company’s history of SQL queries.
In enterprise data systems, the problem is often not a lack of information. It is an excess of it. A warehouse can contain thousands of tables and years of old queries. The model cannot inspect all of that history for every new request, so a retrieval system has to decide which fragments of prior work belong in the prompt.
In Beyond the Harness, posted August 24 with a COLM 2026 workshop journal reference, Kate Gwimm and Carson Eisenach use historical production SQL to build compact reference cards for database tables. Their internal corpus contains 5,176 production queries from what the paper calls a major online retailer. The query being evaluated is excluded from the history available at inference.
The paper’s strongest internal result is structural rather than executable. On a 517-query sample, replacing baseline table documentation with optimized reference cards increases SQL abstract-syntax-tree similarity by roughly 12 percent relative for Claude Sonnet 4.6 and roughly 25 percent for Qwen Coder 3-30B. Optimizing the retrieval-and-prompting harness instead produces roughly 3 percent and 12 percent relative gains. A separate 102-query executable cohort moves in the same direction, but its confidence intervals overlap, so the authors treat those execution results as directional.
The public check is weaker. On a held-out 300-question subset of the BEAVER benchmark, optimized cards alone score 6.67 percent execution accuracy; raw historical SQL scores 6.33 percent. The best result, 9.00 percent, comes from giving the system both cards and raw SQL. That comparison is not statistically significant at this sample size, and the combined condition receives more total context.
The production paper therefore does not establish that distilled reference cards are generally better than raw history. It shows something narrower and more useful here: a real workload can be mined to change the context artifacts that later queries retrieve, and those artifacts can matter materially in the workload where they were learned.
The external result also exposes the central problem. A route learned from one pattern of work may not travel very far when the pattern changes.
Maybe this is just a good index
There is a simpler explanation for all of this: perhaps nothing conceptually new is happening. These are caches, indexes, and ranking systems doing what caches, indexes, and ranking systems have always done.
That rival deserves most of the credit it asks for.
External learning around a fixed model is not new. Reflexion stored linguistic feedback in episodic memory in 2023. A-MEM later built and revised links among memories. HippoRAG 2 showed that a graph-oriented external memory can outperform standard retrieval without learning its structure from the same kind of repeated workload described above. STACKFEED showed that an external knowledge base can be edited from feedback without retraining the base model.
The new papers do not create a third category of learning. They make one engineering boundary easier to see.
The durable record and the route through that record can evolve separately.
That boundary matters even if every component has an old name. A cache is useful because something repeats. An index is useful because some relationships persist. A learned ranking is useful because earlier relevance judgments contain structure that later queries can reuse. None of those benefits is automatic. The question is always what kind of recurrence the system is betting on.
The knowledge-base experiment makes that visible in miniature. Its accuracy gain follows whether relevant keys were previously covered. Its action saving is narrower still, sticking to the exact questions the store had already curated twice. The retrieval experiment relies on stable memory identities and recurring structure across a stream of questions. The SQL result is strongest inside the workload that supplied the historical usage signal and much less decisive outside it.
A learned route is therefore not a general increase in intelligence. It is a wager about what tomorrow will resemble.
A shortcut can become an inheritance
The wager can lose.
A peer-reviewed ACL 2026 study, How Memory Management Impacts LLM Agents, examined agents that retrieve and reuse their own prior executions. The researchers found that when a new task resembles a retrieved old task, the new execution often resembles the old execution too. They call this “experience-following.”
That produces a useful positive-feedback loop when the retrieved example is good. It also produces two documented failures. Errors in old experiences can propagate into later attempts. And an earlier execution that was successful in its own setting can still be a misleading example for a new task. The authors call the second failure “misaligned experience replay.” Their controlled experiments show that regulating the quality of what enters and leaves memory improves long-run behavior compared with naive memory growth.
This is the negative image of the learned-access idea. Prior work can leave behind a shortcut. The shortcut can save effort, but it can also keep steering later work toward the wrong precedent.
That suggests an important asymmetry in how persistent agent systems should be designed.
The underlying record—what was asked, what was tried, what passed, what changed—has archival value. Rewriting it to make retrieval easier can destroy evidence. The derived access layer has a different job. Its links, summaries, rankings, and shortcuts are interpretations of the record. They should be allowed to change because the pattern of work can change.
This is an inference from the evidence, not a result any one paper has established as a universal architecture. But it follows from treating the two kinds of state separately: if a ranking becomes stale, the system should be able to rebuild the ranking without rewriting the events it ranked.
What a benchmark may accidentally measure
Once retrieval experience becomes durable state, another consequence follows.
Two deployments can use the same model and the same source repository yet behave differently because one has inherited a better route through local history. Conversely, a model comparison can be muddied if one model is evaluated inside a mature, workload-shaped retrieval system and another starts from a blank one.
The current papers do not quantify that confound for coding agents. A clean test would have to hold the model, tasks, prompts, tools, and grading constant while varying only the state around retrieval: current files, raw history, a static map of relationships, and a map trained from earlier verified work. It would also need tasks that deliberately depart from the history, so a cache could not masquerade as generalization.
No such software-development experiment has yet established the effect. That is the missing test, not a result to write into the conclusion.
What the existing evidence supports is smaller: in several systems, useful adaptation lives outside the model and outside the raw memory itself. It lives in the changing path between a request and the evidence the model eventually sees.
That path has its own training data, its own costs, and its own failure modes.
The map on the bench
Consider the document store from the opening again. After training, the source documents are still there. What the agent mostly added was navigation: index documents and links pointing into the same underlying material.
On some questions, that navigation made the unchanged reader faster. On new questions whose relevant keys had been covered, it made the reader more accurate. On questions just one step away from the training set, the largest speed advantage vanished.
The map had learned something real. It had not learned everything.
That is a better expectation for persistent agent systems than the idea that more memory simply accumulates into more capability. Repeated work can produce a second artifact alongside the work itself: a route back through what has been learned before.
And when that route eventually points the wrong way, the most valuable property of the system may be that the archive is still sitting underneath it, unchanged, waiting for a better map.
Sources
- Yu Pan and Hongfeng Yu, “Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores”, arXiv, August 22, 2026; submitted to IEEE BigData 2026.
- Qi Feng, Chris Ding, and Jicong Fan, “The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory”, arXiv, August 24, 2026.
- Kate Gwimm and Carson Eisenach, “Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL”, arXiv, August 24, 2026; journal reference: COLM 2026 Workshop — Context Beyond the Window.
- Zidi Xiong et al., “How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior”, ACL 2026.
- Bernal Jiménez Gutiérrez et al., “From RAG to Memory: Non-Parametric Continual Learning for Large Language Models”, ICML 2025.
- Wujiang Xu et al., “A-MEM: Agentic Memory for LLM Agents”, NeurIPS 2025.
- Shashank Kirtania et al., “STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with Feedback”, EMNLP 2025 Industry Track.
- Noah Shinn et al., “Reflexion: Language Agents with Verbal Reinforcement Learning”, NeurIPS 2023.