Using the Past Without Going Back
Agent memory should be evaluated by what historical information changes a later decision, not by direct return alone. TraceML cleanly separates revisitation from score recovery, but its return detector cannot establish whether an agent used history through another path or whether direct return itself caused better outcomes.
A software agent can have a memory and still look, by one reasonable measure, as though it never goes back.
That is the complication hiding inside a new study of machine-learning development. The researchers reconstructed public Kaggle work histories and compared them with agent runs on the same competitions. Their most striking result concerns a behavior they call a solution revisit: after moving away from an earlier version, did the work later become more like that old version than the version immediately before it?
In the study’s twelve-hour comparison, the answer was common enough among the strongest human competitors and almost absent among the agents. The top human cohort produced 531 revisits among 5,838 eligible versions, or 9.1 percent. Codex produced one among 658. MLEvolve produced none among 344.
At first glance, this looks like a memory gap. The paper itself uses that language.
Then comes the awkward detail: MLEvolve already has memory.
Its designers built it to keep several candidate solutions alive, pass information across branches, and retrieve task-specific experience from a Retrospective Memory. The system can therefore preserve and reuse information from another branch without making its current code look like an earlier state on the same path.
The zero is real. What the zero means is not as simple.
The detector sees a return, not every use of history
The distinction matters because TraceML is doing something unusually valuable: it looks at the path of work rather than only the final score.
The full dataset contains 4,465 public human Kaggle trajectories across 134 competitions. Seven competitions form the human-agent paired set. For the behavioral results, the authors narrow the agent side further to a common twelve-hour scope: 10 Codex runs and 107 root-to-leaf MLEvolve branches derived from three MLEvolve runs. The human trajectories are not time- or compute-matched to those agents, so the authors explicitly treat them as a reference distribution rather than a control group.
The revisit measure is also more specific than the ordinary word memory. A later code state has to resemble an earlier, non-adjacent state sufficiently strongly, resemble it more than the immediate predecessor, and have passed through something sufficiently different in between. A plateau does not count. Neither does a useful idea copied from another branch if the resulting code remains different.
That specificity is a strength. It makes the behavior measurable.
It also creates the central boundary of the result. TraceML turns each MLEvolve root-to-leaf path into a trajectory while keeping cross-branch reuse as separate reference edges. If one branch borrows a useful model choice, feature, or lesson from another branch and then continues forward, the history may have affected the work without creating a direct revisit.
So several things that are easy to collapse into one word are actually different:
- an old state can still exist;
- the system can retrieve information from it;
- that information can influence a new path;
- the active path can directly resemble the old state again;
- and the final score can recover after a setback.
A benchmark can measure one of these without measuring the others.
Recovery proves the separation
Codex makes the distinction visible from the opposite direction.
In the same twelve-hour analysis, Codex recovered from 89 percent of measured score setbacks, compared with 79.4 percent for the top human cohort. Yet it produced only one direct revisit under the study’s detector.
A score, in other words, can go down and come back without the work going back.
That matters because the visible outcome is otherwise ambiguous. A recovered score might come from repairing the current approach, generating a fresh alternative, retuning the same system, borrowing something from elsewhere, or restoring old work. Those routes can end at the same number while implying very different capabilities.
The human data does not settle which route is best. Of the top cohort’s detected revisits, 78.5 percent finished above the version they returned to. That is suggestive, but it is not a causal comparison. The study cannot tell us what the same developer would have achieved by continuing forward instead. Nor does a similarity-based detector reveal why a person returned. Public Kaggle versions omit unsaved local work, and the authors infer intent from code rather than observing the developer’s thoughts.
The safest conclusion is therefore narrower than “experts know when to go back.” Humans in this dataset directly revisit earlier code states much more often than the two measured agent systems do. Whether that behavior itself causes better results remains open.
Real programmers go backward, but usually not very far
An older software-engineering study provides a useful independent check on the behavior itself.
YoungSeok Yoon and Brad Myers analyzed 1,459.9 hours of fine-grained Java editing from 21 programmers and detected 15,095 backtracking instances. In 20.4 percent of the cases, the programmer changed code, ran the application or a unit test, and later reverted the change. In 9.5 percent, the reversion was selective enough that conventional linear undo could not express it cleanly.
So technical work really does contain a great deal of going backward. That part is not peculiar to Kaggle.
But the same study supplies an important limit. Ninety-six point seven percent of the detected reversions occurred within the same Eclipse session, and 99 percent within three sessions. The common behavior was local backtracking, not the resurrection of a strategy abandoned weeks earlier.
That makes long-lived branch history a plausible resource, not an automatically important one. A repository can preserve ten thousand old choices. The mere existence of those choices does not show that consulting them improves the next decision.
Sometimes literal rewind really does help
There is, however, evidence that restoring an old state can be useful when the problem is defined differently.
AgentRewind records aligned checkpoints of an agent’s context and controlled workspace. When the agent decides its current trajectory can no longer make progress, it can restore an earlier checkpoint and carry a summary of the failed attempt into the new continuation. On the authors’ 82-task engineering benchmark, the system improves task success and checklist progress over the compared execution strategies across multiple base models and agent harnesses.
That is a real form of going back. It is also a different one from TraceML’s human revisits. AgentRewind begins with a failure and asks where to recover. It does not test whether an agent can recognize that an old strategy—one that may have been rationally rejected at the time—has become valuable because some later fact changed.
Its recovery boundary is concrete too: the runtime can restore files in the controlled workspace, but not external effects such as network requests or outside service state. Even a literal rewind is only as complete as the state the system controls.
This is the strongest case for not dismissing direct return merely because other forms of historical reuse exist. In some failure modes, restoring a coherent earlier state is exactly the operation that matters.
The remaining question is when.
More branches are not the answer by themselves
It would be tempting to turn the TraceML result into a design prescription: keep more alternatives and make agents revisit them more often.
Other machine-learning-agent research argues against that shortcut.
A NeurIPS 2025 study systematically varied both search policies and the operators that modify candidate solutions. Its best pairing raised MLE-bench Lite medal success from 39.6 percent to 47.7 percent, but the important finding was the interaction: search strategy could not be judged separately from the quality of the operation doing the search.
A newer preprint pushes the boundary further. Reasoning as Gradient compares ten models and reports a crossover: tree search retains advantages with weaker reasoning systems, while a more directed optimization method increasingly outperforms as reasoning capability improves.
Neither study measures historical revisitation. Together they make a narrower point: more branching and more exploration are not universal improvements. The value of an old alternative depends on the system that chooses, modifies, and evaluates it.
That is why a useful memory test cannot simply count how much history exists.
Make the past become useful on purpose
A cleaner experiment would force the competing interpretations apart.
Give an agent two plausible ways to solve a software task. Early evidence should make approach A clearly inferior, so abandoning it is the correct decision. Later, introduce a controlled change—a requirement disappears, a dependency behaves differently, a test reveals a new constraint—that makes A, or a small variation of A, the better option.
Then vary only the form of history the agent receives.
One condition gets no explicit old alternative. Another preserves the branch. A third preserves the branch plus the evidence and assumptions that caused it to be rejected. A fourth actively compares new evidence with those old rejection conditions and surfaces the branch when they no longer hold.
The benchmark also needs cases where nothing changes, so a system that compulsively drags old work back into view pays a price for doing so.
The interesting outcome is not the return count. It is whether historical information improves the choice under matched model, tools, budget, and evaluation. An agent could win by restoring the branch, by borrowing one idea from it, or by independently generating an equivalent solution. Those should not be scored as the same internal behavior merely because the final patch is good—and they should not be treated as failures merely because only one looks like “going back.”
That experiment could tell us whether persistence is enough, whether cross-branch reuse is enough, whether explicit relevance cues add value, or whether fresh generation makes old branches largely unnecessary.
Right now, the evidence does not choose among those answers.
The old branch is still there
Version control can preserve a rejected implementation exactly as it was. Months later, the files, commits, and tests may still be available even though the active project has moved far away.
The historical state is durable. Its meaning is not.
A future system might restore that branch. It might retrieve one lesson from it. It might continue forward and never touch it. The useful question is not whether its new code looks enough like the old code to count as a return.
The useful question is whether the past changed a decision that should have changed.
That is a harder thing to measure. It is also much closer to what we mean when we say a work system has learned to use its history.
Sources
- Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, and Yiming Yang, “TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development”, arXiv v1, submitted August 26, 2026.
- Shangheng Du et al., “MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery”, arXiv, submitted June 4, 2026.
- Yu Zhuang et al., “AgentRewind: Recoverable Execution for Long-Horizon LLM Agents”, arXiv, submitted August 14, 2026.
- Edan Toledo et al., “AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench”, NeurIPS 2025.
- Yifei Zhang et al., “Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search”, arXiv, submitted March 2, 2026.
- YoungSeok Yoon and Brad A. Myers, “A Longitudinal Study of Programmers’ Backtracking”, IEEE VL/HCC 2014.