Skip to content

Giving an Agent's Dead Ends Another Job

A completed agent search can become a low-cost test bench for improving orchestration over its recorded branches, giving rejected attempts option value without turning historical replay into a faithful simulator of unobserved futures.

A long search usually has one winner and a great deal of debris.

A coding agent tries one implementation, measures it, branches, backs up, tries another, and eventually finds something good enough to keep. The natural accounting is familiar: the final artifact has value; the failed attempts were the cost of getting there. Save the patch. Maybe save the transcript for audit. Move on.

A new system called Dream-RSI asks what happens if you keep something more structured than a transcript. During an agent-driven search, it records a tree. Each node contains the workspace from which an attempt began, the observations available there, the resulting artifact, the evaluator’s diagnosis, and the score. The parent-child links preserve which attempt grew out of which earlier state.

Then, after the expensive search is over, the system runs a different kind of experiment on the tree.

It does not ask the coding agent to regenerate anything. It does not rerun the evaluator. Instead, it lets alternative management policies decide which recorded branch to reveal next, how many branches to pursue at once, how deep to continue one line, and when to stop. Because the outcomes have already been paid for, thousands of variations on those management decisions can be scored cheaply.

The search history has acquired a second job. It is no longer only a record of work. It is a test bench for deciding how later work should be allocated.

That is a useful idea precisely because it is less magical than it first sounds.

Rerun the manager, not the worker

Dream-RSI separates two pieces of an agent search that are easy to blur together.

The worker generates and evaluates candidate solutions. The orchestration policy decides where that worker should spend effort: open another branch, continue an existing one, run several attempts in parallel, or stop.

During live search, both matter. If a policy chooses a branch, the worker starts from the saved workspace at that branch and produces a new candidate. The paper is explicit that this live transition is stochastic: the same starting state can produce a different child on another run.

Historical replay works differently. When a replaying policy selects a node, the system simply reveals the child that was recorded there during the original search. No new candidate is generated. A different policy can therefore expose a different subset of the same tree, in a different order, and stop at a different point. But it cannot create a branch the original search never produced.

That distinction sounds technical until you notice what it buys. Improving an orchestration policy online is expensive because you have to run the entire search policy long enough to discover whether its choices were good. Historical replay turns those delayed management decisions into cheap feedback. A policy can be revised, tried against the stored tree, revised again, and compared with its predecessor without paying for another round of coding-agent work each time.

The paper then redeploys the best replay-scoring policy into a fresh live search. That new search produces another tree, which becomes another replay world for the next round.

The recursive part of the system is therefore not a model rewriting its own weights. The underlying discovery agent and evaluator stay fixed. What changes is the code that decides how to spend the worker.

This matters because orchestration is not a decorative layer. Once a search can consume hundreds or thousands of model calls, deciding where not to spend them becomes part of the capability.

The clean comparison is smaller than the headline

Dream-RSI’s most dramatic cost claim compares one of its Lasso-search runs with another system that used a different underlying model and tens of thousands of generations. That comparison is eye-catching, but it does not isolate what historical replay contributed.

The more revealing experiment keeps much more fixed.

For its primary controlled baseline, the researchers use the same discovery agent, evaluator, initialization, resource constraints, and initial exploration policy. In the first round, the fixed system and Dream-RSI behave identically. After that, the baseline keeps its orchestration policy unchanged while Dream-RSI revises its policy using the accumulated replay trees.

With Gemini 3.1 Pro on the Lasso task, the fixed policy used 550 discovery-agent calls and produced a solver averaging 3,587.1 milliseconds across six held-out datasets. The replay-improved policy used 317 calls and reached 2,931.0 milliseconds. With Gemini 3.7 Flash, calls fell from 3,200 to 1,879 while average held-out runtime improved from 2,516.7 to 2,350.6 milliseconds.

That is roughly a 1.7-fold reduction in discovery calls in both matched comparisons, accompanied by a better average downstream result.

Even here, the result is not “the replay policy wins everywhere.” The 3.1 Pro solver’s average advantage is heavily influenced by the large RCV1 dataset; on several smaller held-out datasets the fixed-policy solver is slightly faster. The paper’s mathematical-optimization results are mixed as well. Dream-RSI improves some tasks and ties a strong result on another, but is not best on every benchmark.

The useful conclusion is narrower: the stored trees contained enough signal to improve compute allocation in several tested searches.

That is already interesting. It does not need the largest number on the page.

A recorded tree is not the future

The word simulator can make historical replay sound stronger than it is.

A conventional simulator tries to tell you what would happen if you took an action in some state. Dream-RSI’s replay tree can tell you what did happen once, when that action was followed during the recorded search. If the live worker is stochastic, those are different claims.

The paper’s implementation makes the boundary unusually clear. Online, selecting a branch causes a new generation and evaluation, and the same starting state may yield a different result next time. Offline, selecting that branch returns the recorded child deterministically. The replay policy remains inside the support of the historical tree; no outcome beyond that tree is generated.

So if a replay policy learns that branch B was excellent and branch C was terrible, it has learned something true about the stored search. Whether that ranking survives a new rollout depends on how representative those sampled outcomes were.

The system does include an important safeguard against one trivial form of regression. The current orchestration policy is always included among the candidates, so the selected successor cannot score worse than the incumbent on average across the fixed replay histories. That is a guarantee about replay score, not about the next stochastic live search.

This is the same general problem that makes counterfactual reasoning from agent logs treacherous. A separate 2026 study, The Replay Gap, examined a different intervention—switching the model inside a running software agent—and found that the substituted model quickly changed later actions, invalidating a static replay of the original trajectory. Across its experiments, model swaps rewrote 61 to 94 percent of post-fork actions.

That study does not refute Dream-RSI. Dream-RSI keeps the worker fixed and changes the policy that chooses among recorded branches; The Replay Gap changes the worker inside a closed loop. But the comparison exposes the same discipline: a log is authoritative about the world that was observed. It is not automatically authoritative about the world an intervention would have created.

For Dream-RSI, the missing experiment is therefore not another replay score. It is repeated live continuation from matched saved states. If a branch that looks promising in the frozen tree also tends to produce good fresh children when rerun, replay is learning a transferable allocation rule. If its advantage vanishes under reruns, the policy may be learning the luck of a particular tree.

The paper does show successful fresh redeployment after replay optimization, which is why this is more than an offline curiosity. But it does not report the kind of repeated-run uncertainty that would tell us how stable those gains are under the stochastic transitions it describes.

The failed branch changes category

Once the boundary is clear, the practical consequence becomes more interesting.

Suppose an agent opens ten branches and only one produces the artifact that ships. For the current task, the other nine may still be failures. They consumed compute and did not become output.

But if each failed branch retains its parent state, cost, evaluator result, and relation to neighboring attempts, it can become evidence about a different question: how should the next search be managed?

One branch may show that a promising direction required too many refinements before paying off. Another may show that late parallelism mostly produced redundant candidates. A third may reveal that opening a fresh branch early was cheaper than repairing a deeply explored one. None of those facts makes the failed artifact successful. They give the failure option value as orchestration data.

This is different from turning transcripts into training examples or asking a model to summarize “what we learned.” The tree is useful because it preserves decisions and consequences in a form another policy can act against. The manager can be tested on the actual branching structure instead of merely being told a lesson in prose.

That distinction also explains why a pile of logs is not enough. If history loses the workspace state, branch parentage, evaluation outcome, or timing of what was visible when a decision was made, replay becomes ambiguous. A transcript may document that two attempts happened. A usable replay record needs to say which state each attempt inherited and what happened after the policy chose it.

The object worth preserving is therefore not simply “agent memory.” It is decision-grade lineage.

History helps only when it transfers

There is a temptation to turn this into a general rule: save everything because failed search will train tomorrow’s manager.

The evidence does not support that.

A frozen tree is an irregular sample of one search. Some branches are dense; other possibilities were never tried. A replay policy can only learn from the recorded support. It may also exploit quirks in the evaluator, the task family, or the particular random outcomes that happened to appear in its trees.

And adaptive orchestration itself is not new. Systems such as the Darwin Gödel Machine already treat the process that generates and evaluates candidate agents as something that can improve over time. The novelty in Dream-RSI is more specific: completed discovery histories are reused directly as cheap empirical worlds for testing executable orchestration policies, while the expensive worker stays idle.

That gives the idea a natural kill test. Optimize a management policy on recorded trees, then evaluate it on unseen tasks and repeated fresh continuations from the same saved states. Compare it not only with a fixed policy, but with simpler adaptive rules: stop when recent improvements flatten, allocate by a bandit heuristic, or cap unproductive branches. If the replay-trained policy cannot beat those controls outside the trees that taught it, the extra machinery is mostly an elaborate way to memorize old luck.

If it does transfer, the accounting of agent work changes.

The value of a search would no longer end with the artifact it produced. Some of the expensive exploration could become infrastructure for later allocation decisions. A branch that failed to ship might still reduce the cost of deciding where the next hundred model calls should go.

That does not make waste disappear. Most dead ends can remain dead ends. Historical replay simply creates a second market in which a recorded attempt can be valuable: not as an answer, but as evidence about what to try next.

The important word is recorded. A failed branch forgotten by the system is still gone. A failed branch preserved with enough lineage can be asked a new question.

It still cannot tell you what would have happened on a road never taken. But it may tell you that, next time, you should spend less time on the roads you already know.

Sources