Skip to content

What Should Survive a Coding Agent Handoff?

Related:Cards ↗

A model handoff should treat durable work state and receiver-visible history as separate design variables: preserving a complete record is not the same decision as using that record to condition the next model.

A coding agent stops halfway through a repair. Its edits are still on disk. The failing tests are still there. The repository remembers what changed.

Now another model takes over.

There is a second kind of state that may cross that boundary: the conversation. The successor can inherit every hypothesis, command, tool result, mistake, and explanation produced by the first model. Or it can inherit much less while starting from the same kind of persistent work state.

It is tempting to call the first option continuity and the second amnesia. A new coding-agent study shows why that distinction is too simple.

The Handoff Tax paired a cheaper, lower-capability model with a costlier, higher-capability model in each of the Claude and GPT families and ran the combinations across all 500 SWE-bench Verified tasks. The study switched models at seven points in a run and compared four continuation interfaces: the full trajectory, two kinds of summary, or no predecessor trajectory. In every interface, edits already written to the working tree remained available to the successor.

When the lower-capability model handed upward, removing its trajectory improved the higher-capability successor’s aggregate pass rate. For Claude, the reported average rose from 69.2 to 72.4 percent; for GPT, from 67.5 to 79.7 percent. When the direction reversed, removing the trajectory hurt instead: Claude fell from 65.6 to 60.9 percent and GPT from 81.0 to 75.4 percent. Task-clustered bootstrap intervals for all four trajectory-removal-versus-raw differences excluded zero.

Those percentages need a boundary around them. They are unweighted means across seven fixed switch points, with each point evaluated on the intersection of tasks that actually reached a handoff under all four interfaces. The coding study also ran one episode per task under each configuration, so it does not estimate repeated-run variability for the same task and interface. The result is evidence about these two model pairs in this benchmark, not a universal law of escalation.

Still, it creates a useful anomaly. The value of old context changed when the receiver changed.

The obvious lesson is wrong

The easy takeaway would be: when a stronger model arrives, start it clean.

An independent takeover study makes that rule hard to defend.

Handoff Debt began with 75 SWE-bench Verified tasks and generated 181 interruption points. At each point, the researchers froze the repository checkpoint and let a successor continue with one of four views of the predecessor: the repository alone, a raw trace, free-form summary notes, or structured notes.

Across three successor models, context-bearing handoffs reduced median agent events by 20 to 59 percent and cumulative prompt tokens by 42 to 63 percent compared with repository-only takeover. Raw traces cut events by 57 to 59 percent and prompt tokens by 50 to 59 percent. They also raised observed solved rates by 6.1, 6.6, and 14.9 percentage points, although the paper treats those solved-rate effects as smaller and more model-dependent than the efficiency result.

The missing transcript was not harmless. It forced the successor to rediscover work that the predecessor had already done.

So these studies are not really arguing for opposite policies. They are exposing two different jobs that the same transcript is asked to do.

A transcript is evidence and a path

Consider what sits inside a long coding trajectory.

A test result is an observation. A command that exposed a dependency may be difficult to reconstruct. A list of files already inspected can save time. Those are reasons to carry history forward.

But the same trajectory also contains guesses. It records which explanation became salient, which file the first model obsessed over, which failed approach occupied twenty turns, and which assumptions were repeated until they started to look settled. Those are not the same kind of state as a test result.

The Handoff Tax experiment does not isolate which of these ingredients caused its directional reversal. Removing the trajectory removes useful observations, bad hypotheses, tool output, verbosity, and token load at the same time. Calling the effect “anchoring” would outrun the experiment.

There is, however, independent evidence that model-authored history can influence later reasoning even when relevant evidence remains available.

In Contextual Drag, Princeton researchers put failed draft solutions into the context of 11 proprietary and open-weight models across eight reasoning tasks. They report performance drops of roughly 10 to 20 percent, and later reasoning often became structurally more similar to the failed draft. External error feedback and successful self-verification did not fully eliminate the effect. That is not a coding-handoff replication, but it shows that a model can recognize bad prior work without becoming independent of it.

A second independent study looked at model switching directly. Researchers at NatWest AI Research and University College London ran a nine-by-nine matrix in which one model wrote the early turns and another produced only the final turn. On CoQA, 16 of 72 cross-model switches were statistically significant at the 95 percent level; on Multi-IF, 18 of 72 were. The effects were directional: a prefix that hurt one receiver could help another.

CoQA provides the cleaner diagnostic. The original passage remained in the dialogue history. The new model therefore still had access to the source text, yet some receivers stayed consistent with earlier assistant answers rather than fully re-grounding on the passage. The source of truth was present. The inherited conversational state still mattered.

Taken together, the studies support a narrower explanation than “bad reasoning poisons good models.” A trajectory is a mixed object. It can carry evidence that saves rediscovery, and it can carry a predecessor’s path through that evidence. Different receivers can value that mixture differently.

The task can reverse the answer

The Handoff Tax paper contains its own strongest warning against turning the coding result into a rule.

Its four-interface comparison was established in SWE-bench. The authors also ran raw handoffs in two settings with different information patterns.

In 535 Lost in Conversation examples, requirements arrived over several turns. The receiver was the first model that could see the fully specified task. Under that setup, raw escalation recovered 86 percent of the higher-capability model’s quality advantage in the reported Claude aggregate.

In a separate BrowseComp experiment, the question was available from the beginning but useful evidence accumulated through web search. On 200 browsing-required questions, GPT raw escalation recovered 95.8 percent of the higher-capability model’s quality advantage. Inheriting the earlier search also shortened the higher-capability continuation by about three calls on average.

This suggests a more useful question than “how much context should the next model get?”

What is in the context that the receiver would otherwise have to rediscover?

If the predecessor has already found scarce evidence, history may be valuable. If the decisive requirement only arrives after the switch, the successor may be relatively free to reinterpret the problem. If the entire task was visible from the beginning and the predecessor spent many steps inside an unproductive theory, the old trajectory may carry more path than evidence.

Information has a history too. When it became available changes what a handoff can profitably preserve.

Routing has a second decision

Model routing is usually framed as a choice about who works next. Start cheap. Escalate when the task looks hard. Downshift after the difficult part is over.

Long-running work adds another decision: what does the next model inherit?

The two decisions need not share an answer. A cheap model’s trajectory may be useful evidence that an escalation is necessary without being the best prompt for the model that receives the escalation. Conversely, a weaker successor may need the stronger predecessor’s discoveries to avoid repeating expensive work.

That leads to an architectural distinction the experiments do not themselves prove, but do make worth testing: preserve the audit record separately from the receiver’s working context.

A system could retain the complete trajectory for debugging, review, provenance, or later reconstruction while exposing the next model to a deliberately constructed handoff surface. That surface might contain the current task, changed files, verified test results, unresolved facts, and rollback information without automatically replaying every predecessor hypothesis.

The strongest version of that proposal is still untested. A fair experiment would freeze one exact repository checkpoint and compare several receiver surfaces: repository only; the full raw trace; a compact bundle of verified observations and failed tests; a predecessor-written rationale; and a length-matched block of neutral text to measure context burden by itself. Repeating each condition would separate interface effects from run-to-run stochasticity. Crossing several ordered model pairs and task types would test whether the direction seen in The Handoff Tax generalizes at all.

That design can fail in informative ways. If neutral context causes the same damage as failed reasoning, token load may explain much of the apparent path dependence. If raw traces win once routing is adaptive rather than fixed, selective exposure may matter less in production. If a verified-state bundle preserves the rediscovery gains of Handoff Debt without reproducing the coding penalty, the case for separating record from prompt becomes stronger.

Goodfoot publicly describes Cards as a control surface for directing and reviewing coding-agent work. That makes it a plausible place to run such a test, not evidence for the thesis. The experiment would require instrumentation that can retain the full history outside the receiver’s prompt, hold the work state fixed, and vary only what is exposed at takeover. The current public product description does not establish that Cards already supplies all of those pieces.

The record is not the prompt

Return to the interrupted repair.

The first model’s edits do not have to disappear when the second model arrives. Neither does its transcript. The useful distinction is between retaining that transcript and making it part of the next inference.

One decision asks whether the organization can reconstruct what happened. The other asks what should shape the successor’s next search.

The current evidence does not tell us the ideal handoff format. It does tell us that those questions should not be collapsed into a single setting called “context.”

A future handoff might preserve the complete journey in the record while giving the next model only the task, the current work, and the pieces of history that have earned their place. The model can still inspect the rest if the system makes it available. But history would become a controlled input rather than an automatic inheritance.

The work survives. The record survives. What changes is the assumption that the next model must relive both in order to continue.

Sources