Skip to content

When Yesterday's Fix Sends an Agent Astray

On selected coding tasks in VibeMemBench, failures often occurred in the records memory systems supplied after reaching relevant history; construction, context burden, and downstream use therefore need outcome checks beyond retrieval, with benefits bounded by the recipient and task.

A coding agent could already complete a particular aiohttp repair in all four of its test runs. Then a memory system handed it 334 lines from an earlier task.

With that record in context, the agent failed all four runs. A different memory system supplied one line for the same target, and the same solver remained at four successes out of four.

The comparison appears in VibeMemBench, a September 2026 benchmark for memory systems in repository coding. It is a small, selected case, and the two systems supplied different content as well as different lengths. The researchers had not simply compressed one identical record and proved that shorter was better.

But the longer record left a useful trail. According to the paper’s analysis, it described another issue’s report, reproduction steps and instructions. The solver was redirected toward an exception-handling block rather than the behavior it had repaired successfully without memory.

The system had retrieved history. The history had changed the work. Those two facts did not make the retrieval helpful.

Memory systems promise to save agents from rediscovering what earlier work has already established. That promise requires more than finding a related episode. Someone or something must decide which part of the episode to preserve, what form to give it and how it should influence a new task. A record can point to the right neighborhood while carrying the wrong repair, or bury a useful lesson inside material the worker does not need.

VibeMemBench follows what happens after that handoff. Its most revealing question is not whether the past can be found. It is what the memory system turns the past into.

A stored coding trajectory contains several kinds of material: the original request, file names, commands, tool output, tentative explanations, failed edits and perhaps a successful patch. Retrieving it does not decide which of those things should become advice.

VibeMemBench makes the eventual repair executable. The agent receives a repository task and a declared memory condition, then edits the code. Tests determine whether it succeeds. The researchers compare runs with and without the supplied record under the same agent, tools, sandbox and budget.

They also make an unusual selection decision. From 3,634 historical trajectories across 90 repositories, they retain 111 targets for which a chosen historical experience had already improved a reference solver. This gives them a way to ask whether a memory system can deliver something useful from that history.

It also limits the claim. These targets were selected because a record helped in one reference setting, with only four seeds used to choose among candidates. They do not reveal how often arbitrary repository history contains a transferable lesson. When the researchers froze the selected experiences and gave them directly to five other solvers, four showed small positive average changes, but every confidence interval included zero.

The separate memory-system test was less encouraging. Mem0, SimpleMem, MemoryOS and A-MEM constructed records from the eligible history and retrieved one for each new task. In 11 of the 12 solver-system pairings, the observed success rate was at or below the matched memory-off baseline. None had a confidence interval entirely above zero; one was significantly worse under the reported bootstrap.

Those results leave several possible explanations. A system might fail to store the relevant episode, retrieve the wrong record, lose the useful detail while constructing it, or supply good material that the solver mishandles. A failed repair alone does not distinguish them.

The authors’ audit of 231 failing pairings points toward the supplied records. Under their scripted labeling rule, pure ranking misses were rare. More often, related history had surfaced but its repair lesson was missing, generalized too far or buried in raw transcript.

Two blind model judges were less consistent about exactly where ranking failure ended and poor record construction began. The stable conclusion is therefore coarser than the taxonomy: many failures concerned the material handed to the solver, not simply whether the archive contained a relevant episode.

The two styles of memory could fail differently. Transcript-based records retained old dialogue and tool noise. Compact records could lose the conditions that made a repair valid, or reduce it to a generic description of the symptom. Finding something related did not settle what should be said about it.

Were the old instructions the problem?

The 334-line record suggests an appealing explanation: old instructions had been mistaken for current instructions. Remove them and the agent should recover.

The researchers tested that possibility. For 62 failing pairings flagged for instruction pollution, they removed the flagged lines and supplied the remaining record again. Performance improved.

Then they tried a more ordinary intervention: remove the same number of randomly chosen, unflagged lines.

That improved performance almost as much. The study’s predeclared criterion for a special instruction-content effect failed. The result weakened the explanation that the instruction-like wording itself was the cause and favored the simpler possibility that these records imposed too much transcript burden.

The distinction changes the remedy one might expect to work. A filter searching for imperative language could remove conspicuous commands while leaving the worker to process an unnecessarily large record. Shortening can help without the deleted lines having been malicious or authoritative.

Independent evidence makes that possibility credible. In Context Length Alone Hurts LLM Performance Despite Perfect Retrieval, researchers varied input length while keeping relevant evidence retrievable. Performance still declined in tested reasoning and coding settings; whitespace and masking controls showed that ordinary retrieval failure was not a complete explanation.

That study is not a replication of the coding-memory experiment. Nor does either result establish that every long record should be shortened. The deletion test changes the material supplied as well as its volume. What it establishes more securely is that spotting instruction-like text was not enough to identify the harmful mechanism.

A memory system may have to preserve less, but it still has to preserve the right thing.

One line can carry the wrong lesson

A different aiohttp case prevents the story from becoming a rule about length.

Here, SimpleMem supplied a one-line record describing an old edit that was itself wrong. It contained useful file and function names alongside a misleading repair direction.

One solver used the names as landmarks and repaired the parser’s state machine correctly. Its observed success count rose from one run out of four to four. Another almost reproduced the erroneous historical edit, falling from two successful runs to zero. The same line had served as a map for one worker and a recipe for the other.

The paper’s trace analysis supports that description. Four runs per condition still leave substantial uncertainty about how stable the difference is. There is also a less psychological explanation for many of the study’s cross-solver reversals: the amount of room each solver had to improve.

Forty-two percent of the memory-system pairings began with four successes out of four without memory. On that bounded count, a record could only tie the result or reduce it. A solver starting at zero could only tie or gain. Those are observed four-run endpoints, not proof that either solver would always succeed or fail on the task.

This headroom pattern explains much of the direction of the reported changes. It cautions against attributing every difference to a hidden faculty called trust, or deciding that a record has a stable positive value because it once helped another model.

It also clarifies the post-retrieval job. The worker must determine which parts of a historical account remain applicable. A file name can be useful even when the edit described beside it should be rejected. Compressing both into one confident sentence can make the record shorter without making that distinction easier.

A record can advise without deciding

A second benchmark examines that final judgment directly, though in a different setting.

MemCalib supplies constructed blocks of memory rather than asking a retrieval system to find them. Its rubric assigns each proposition a role: some should be ignored, some should provide limited support, and some should control an important conclusion or constraint.

Those roles expose two possible errors. A model can let an irrelevant or merely advisory proposition dominate its answer. It can also neglect information that the task says should govern the response. “Use the memory” is not a sufficiently precise instruction when one block contains both kinds of information.

MemCalib evaluates constructed examples across health, general assistance and coding. Its coding tasks ask for written planning, code understanding and diagnosis rather than executable repository patches. Models ran in non-thinking mode, and another model judged their use of the propositions. The resulting measurements are useful evidence about selective use, with uncertainty especially at the boundary between supporting and controlling information. They do not explain the cause of VibeMemBench’s failed repairs.

The broader problem also predates this benchmark. Earlier work on persistent memory and personalization had already found over-application of old preferences and failures to use relevant constraints. MemCalib makes the degree of influence explicit within a composite record.

That is a different responsibility from retrieval. Selecting a record because it resembles the current task does not establish that every sentence in it should receive the same authority.

Memory can still earn its place

The failures do not imply that persistent memory is a bad idea. Other designs report positive executable outcomes.

MemGuard, published a month earlier, keeps verifier-derived quality and status information attached to records as they are admitted, retrieved, combined, summarized and archived. Across four model backbones and four terminal, coding and web benchmarks, its authors report the best success metric and fewest average steps in all 16 combinations they tested.

Those are different tasks and protocols from VibeMemBench. The comparison does not identify which individual memory operation is best, or demonstrate that MemGuard would repair the failures described here. It does prevent a field-wide conclusion that current memory systems cannot help agents.

The more useful distinction is between retrieving an episode and preparing it for reuse. A memory system’s output becomes part of the next worker’s task environment. Its construction decisions can preserve a constraint, erase it, repeat an old mistake or make a small clue compete with hundreds of unnecessary lines.

That places responsibility on the system producing the record as well as the solver reading it. Retrieval can establish a connection to earlier work. The record still needs enough of that work’s conditions for the new worker to judge what applies.

In the one-line aiohttp case, the useful landmarks and the bad edit traveled together. One worker separated them. The other carried the mistake forward.

Sources