Should an Improving Agent Keep Going?
On continuously scored open-ended agent tasks, absolute improvement and marginal compute value can diverge: after a measured crossover, an active session may still improve while additional breadth becomes the better use of the same fixed budget.
The most misleading sign in a long agent run may be a better answer.
If an agent has just improved its best result, continuing feels easy to justify. The run has not crashed. It has not gone in circles. The latest checkpoint is better than the previous one. Whatever the session has accumulated—tool results, failed approaches, feedback, a working theory—appears to be paying rent.
A new study asks a different question: not whether the run is still improving, but whether it is still the best place to spend the next unit of compute.
That distinction produced an odd result. In When Agents Slow Down, researchers tracked agents on open-ended tasks where intermediate submissions receive continuous scores. The best result found by a session could keep getting better even as the session’s marginal gains fell below a reference for simply starting fresh attempts. On one FrontierCS Polyomino Packing experiment, the researchers held the total budget at 100 million cache-inclusive tokens. The measured inflection point was about 38 million tokens, implying a three-session allocation. That allocation finished 264 Elo above one 100-million-token run and 355 Elo above ten 10-million-token runs. The paper’s 90% session-bootstrap intervals for those gains were +135 to +412 and +226 to +618 Elo, respectively.
The winning policy in that experiment was neither “keep going” nor “parallelize everything.”
The useful discovery was smaller and stranger: progress and allocation value are not the same measurement.
A score can rise while its slope loses
Open-ended optimization tasks are unusually good places to see this distinction because they do not wait until the end to say whether anything worked. An agent can submit a packing, a kernel, an algorithm, or another artifact repeatedly and receive a score each time. That creates a curve rather than a single pass/fail result.
The authors of When Agents Slow Down turn those curves into what they call Elo-per-token. At each token budget, they keep the best score a session has found so far. They then compare those best-so-far results within each task and fit a Bradley–Terry rating model, which lets tasks with different native score scales contribute to a common Elo curve.
The important quantity is not the height of that curve. It is the slope.
The paper derives a reference for independent sampling. Under its assumptions—attempt scores are independent and identically distributed draws from a continuous distribution, and attempts have equal cost—multiplying the sampling budget by ten adds 400 Elo in expectation. That gives the adaptive agent a yardstick. Early in a run, a session can convert tokens into improvement faster than that fresh-sampling reference because continuity has real advantages: it can use feedback, refine a promising incumbent, remember what failed, and exploit discoveries it already paid for.
Later, the measured slope can fall below the reference even while the best-so-far curve keeps rising.
Nothing magical happens at that crossing. The agent has not suddenly become incompetent. Its current best answer does not get worse. The paper’s “scaling inflection point” is an opportunity-cost boundary: the point where another chunk of compute spent deepening this session is expected to buy less Elo than the same compute spent on additional breadth under the reference model.
That is why an improving run can become a bad bet.
The Polyomino Packing intervention makes the idea concrete. The measured curve implied an interior split of the fixed budget. The resulting allocation beat both extremes: one marathon session and ten much shorter sessions. The point is not that three sessions are universally optimal. In the paper’s second allocation experiment, on an MLS-Bench load-balancing task, the nearby one-, two-, and three-session allocations were not statistically distinguishable. The point is that a fixed budget can have a depth–breadth optimum in the middle, and the shape of the run can help locate it.
That is a different decision from asking whether more tokens still help.
Continuity really can be valuable
A restart slogan would be a convenient takeaway. It would also conflict with some of the best evidence we have about long-running agents.
EdgeBench was built around 134 tasks that can sustain at least 12 hours of agent work, with feedback from tests, evaluators, simulations, judges, or other task-specific mechanisms. Across roughly 38,000 hours of interaction, its authors report that long-horizon performance depends on accumulated experience, not merely the number of attempts. In one matched 12-hour comparison across 17 tasks, a continuous Opus 4.8 run averaged 43.0 versus 36.1 for the best of six independent two-hour restarts. A separate context-length ablation also favored the longer context throughout the measured window.
That result is not a contradiction to the slowdown paper. It is a warning against reading the crossover as a moral judgment about long sessions.
Continuity has value when history contains information that changes what the agent can do next. A failed experiment can rule out a region of the search space. A test result can expose a hidden constraint. A partial proof can make a later proof cheaper. Throwing that state away can force the next run to pay again for discoveries the first run already earned.
The question is marginal. A session can be worth continuing at ten million tokens and not worth continuing at eighty million. It can be better than restarting from scratch and still be worse than opening a second trajectory from a carefully preserved work state. The value of continuity is not a property that a run either has or lacks. It can change over the life of the run.
This is familiar outside AI. A factory can still be producing salable units while the next dollar belongs on a different machine. A researcher can still be improving an experiment while a second line of inquiry has the higher expected value. “Still working” describes the local process. Allocation asks what the scarce resource should do instead.
Agents make the distinction easy to miss because the scarce resource and the learning process are tangled together inside one conversation.
Breadth has failure modes too
Once a long run loses its marginal edge, it is tempting to imagine that fresh sessions automatically buy independent exploration.
They may not.
The DivInit study starts from the opposite problem: ordinary parallel samples from the same model and prompt can be redundant. The authors deliberately diversified the initial queries given to parallel agents and reported better matched-compute performance across five open-weight models and eight benchmarks. Fresh processes are not necessarily fresh ideas.
That matters for the sampling reference in When Agents Slow Down. The 400-Elo-per-decade line is a theoretical result under independent, identically distributed, equal-cost attempts. Real agent branches can share the same prompt, model, repository, examples, retrieval results, and obvious first move. Ten nominally separate sessions can therefore spend much of their budget rediscovering the same basin.
The Polyomino result hints at exactly this practical tension. One long run had too much depth relative to its late marginal return. Ten short runs had too little depth. An interior allocation gave each trajectory enough time to exploit early learning without committing the entire budget to one history.
Other work arrives at a compatible boundary from a different direction. Benchmark Test-Time Scaling of General LLM Agents reports limits to both sequential scaling and parallel scaling: long sequential trajectories encounter a context ceiling, while parallel attempts encounter a verification problem because generating more candidates is only useful if the system can reliably choose among them.
Breadth, in other words, creates inventory. Someone—or some verifier—has to distinguish the useful branches from the merely different ones.
A compute policy for agents therefore has at least three moving parts: how much a trajectory is still learning, how independent a new branch would actually be, and how expensive it is to evaluate and reconcile the additional candidates. A crossover in token efficiency answers the first question. It does not erase the other two.
The cause of slowdown is still unsettled
The clean part of When Agents Slow Down is the measurement. The causal story is not clean yet.
The authors discuss a “sticky basin” possibility: an early strategy can become difficult to abandon as a session accumulates history around it. That is plausible. It is also not isolated by the experiments.
There is an important harness caveat too. The paper’s public implementation says the 100-million-token headline runs use cache-inclusive accounting and that a watchdog resumes an agent in the same session when it exits with budget remaining. The paper notes that models tend to declare completion well before the target. Some late-run behavior may therefore reflect work forced beyond an agent’s own stopping decision, not only a spontaneous desire to continue.
Several other mechanisms could produce a similar late curve. A long context may preserve too much low-value history. The task itself may have fewer remaining improvements that are easy for this model to discover. The evaluator may become less sensitive to late improvements. A fresh session may help because it changes the reasoning history, or simply because it gets another random draw at the search problem.
Those explanations predict different interventions.
If the problem is inherited reasoning, a new model context starting from the same artifact should recover exploration. If the problem is opportunity depletion in the artifact itself, fresh contexts from that checkpoint should stall together. If a distilled record beats both the full history and a blank restart, then the important variable is not continuity but which state survives it. If deliberately diversified starts beat naive restarts, then “more sessions” was never the real treatment.
The current paper does not settle those tests. Its public code is available, but its README says the raw trial trajectories, logs, and aggregate data will be released separately; as of September 16, that release had not happened. The headline results therefore cannot yet be independently regenerated from the underlying trial data.
That is why the paper’s strongest result is economic rather than psychological: it identifies a point where the observed return on continuation falls below a comparison policy. It does not yet tell us why.
That boundary matters. “The agent got stuck in its context” is a causal claim. “The next tokens had a lower measured return than additional breadth” is an allocation finding.
Only the second has been demonstrated here.
Stop the trajectory, not necessarily the work
Software makes the unresolved distinction especially consequential.
A coding-agent session can contain at least two kinds of accumulated state. There is durable work: files changed, tests added, benchmark results, a Git commit, a reproduced failure. Then there is trajectory state: the sequence of hypotheses, tool calls, explanations, and choices that shaped the current search.
Those states do not have to share a lifetime.
Suppose a performance-tuning agent reaches a repository checkpoint that is clearly better than the starting point, but its measured improvement rate has flattened. The next experiment need not choose between “continue everything” and “throw everything away.” It can preserve the improved repository and fork new reasoning from that exact artifact. Another branch can inherit a distilled set of verified facts. A third can keep the full history. The artifact remains fixed while the trajectory becomes the experimental variable.
That is where the slowdown result becomes relevant to systems such as Goodfoot’s Cards and Upstream: not as evidence that either product solves the problem, but as instrumentation for asking a falsifiable question. If the work state and the agent history are recorded separately, a team can compare continuation against fresh or deliberately diversified trajectories without pretending that a restart means returning to the original code.
The outcome could cut against the thesis. Fresh context may fail to recover anything. Review and merge costs may erase the apparent gain from extra branches. The inflection estimated on earlier runs may fail to predict held-out runs. A natural stopping policy may reallocate compute just as well without profiling a curve first.
Those are not edge cases. They are the tests required before turning one benchmark result into an engineering rule.
What an improving run no longer tells you
The old stopping question is binary: is the agent still making progress?
On the tasks in When Agents Slow Down, that question can return “yes” long after it stops resolving the allocation problem. Best-so-far quality may still rise. The session may still be doing intelligent work. Its history may still contain valuable information.
None of those facts says that the next token belongs there.
The more useful question compares slopes: what is this trajectory buying now, and what could the same budget buy somewhere else?
That comparison also explains why the opposite defaults can both fail. Ending every long session discards useful learning. Letting every improving session run indefinitely ignores opportunity cost. Launching as many parallel attempts as possible can replace one stale trajectory with ten redundant ones.
On Polyomino Packing, the better allocation sat between those instincts. The long run had not stopped producing better answers. It had stopped producing them fast enough to deserve the whole budget.
That is a narrower claim than “agents stop learning.” It is also a more useful one. A better answer can be real progress and still be evidence that the next unit of compute should go somewhere else.
Sources
- Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, and Alvin Cheung, “When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis”, arXiv, September 14, 2026.
- Agent-TTS-Code, the authors’ public implementation for the experiments and Elo analysis.
- Deyao Zhu et al., “EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments”, arXiv, July 6, 2026.
- Xiaochuan Li et al., “Benchmark Test-Time Scaling of General LLM Agents”, arXiv, February 22, 2026.
- Sidhaarth Murali et al., “Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search”, arXiv, June 15, 2026.