Skip to content

How Do Two Agent Workspaces Become One?

Related:Cards ↗

Useful parallelism depends not only on task decomposition but on whether independently produced intermediate states have an acceptable way to rejoin; changing the representation or coordination protocol can change which team structures are feasible.

On one recent computer-use benchmark, adding more agents made a strong system worse.

The single agent completed 55 percent of the tasks. A multi-agent system, built to split work into parallel subtasks, completed 50.5 percent. Then a different multi-agent design reached 67 percent on the same 200-task benchmark.

The natural question is what the third system did better. The answer in Spine-Branch Coordination for Multi-agent Computer Use, posted August 22, is partly about something that workflow diagrams usually hide: what happens to the world in which the work was done.

A browser task does not produce only an answer. It can also produce an authenticated session, open windows, cookies, files, running processes, a half-completed form, local application state, and changes already made to remote services. If two agents begin from the same virtual computer and then take different actions, those environments diverge.

A later worker can inherit one of them. There is no general instruction called “combine these two computer worlds correctly.”

The new system therefore keeps one continuous virtual-machine lineage for the main task. The paper calls that path the spine. Work is sent to parallel branches only when the useful result can be brought back as a self-contained artifact—information, a file, or another result that does not require preserving the branch’s whole machine. When a branch finishes, its artifact survives and its machine can disappear.

Across three computer-use backbones on the Odysseys benchmark, the authors report gains of 6.0 to 16.5 percentage points over the multi-agent baseline and per-task cost reductions of 34 to 70 percent. Those are substantial results. They are not a clean causal experiment on state handling alone. The new system also changes planning, information routing, context management, manager usage, and prompting. The paper shows that a design built around state lineage can work much better; it does not tell us how many of those percentage points came from the lineage rule itself.

That limitation turns out to make the broader lesson more interesting.

The important question is not whether every multi-agent system needs a spine. It is whether a plan for parallel work has specified how the branches are allowed to come back together.

The missing operation

A common way to plan agent work is as a graph. Some tasks depend on others. Tasks without dependencies can run at the same time. When two branches are finished, a later step consumes their outputs.

That is the basic design of Multi-Agent Computer Use, the earlier system used as a comparison in the Spine-Branch paper. Its manager breaks a task into a directed graph, sends ready subtasks to agents in parallel, and revises the graph as new information arrives. In its original experiments, that design improved on strong single-agent baselines across several computer-use benchmarks.

A dependency graph answers one question very well: what has to happen before what?

It can leave another question unstated: what operation turns the outputs of two completed branches into one valid next state?

Sometimes the answer is easy. Two researchers can return notes. Two programs can produce separate files. Two measurements can occupy different rows of a table. The downstream step can receive both without deciding which prior world should survive.

Sometimes the answer is not easy because the output is entangled with a history.

Consider two agents working from cloned browser sessions. One fills half a form. The other changes an account setting. Even if both did useful work, “merge the sessions” is underspecified. Which cookies should survive? Which form values? Which cached data? Which remote action has already happened? If both changed the same record on the server, what would it even mean for a merged desktop to undo that disagreement?

The difficulty is not that the machines cannot be copied. It is that copying is different from reconciling.

The Linux project CRIU can freeze a running application or container, save its state, restore it, and support live migration. That is an impressive way to move one execution history. It still does not supply a universal rule for turning two independently changed histories into the one history a task intended.

The distinction matters well beyond virtual machines.

Some things are designed to rejoin

Software developers already work with a system whose whole purpose is to make divergent histories easier to combine.

Git gives branches ancestry, diffs, and explicit merge behavior. When two histories cannot simply be fast-forwarded, Git can create a merge commit that records both histories as parents. But Git’s own documentation also shows the limit: if both sides changed the same area in ways Git cannot reconcile automatically, the merge stops and leaves the conflict for someone else to resolve.

Distributed systems contain a cleaner example. Some replicated data types are deliberately designed so different replicas can accept updates independently and later converge. Marc Shapiro and colleagues’ 2011 paper on conflict-free replicated data types formalized conditions under which that convergence is guaranteed.

These systems make an important point easy to see. A team does not become easier to coordinate only because its members get smarter. Coordination can become easier because the thing they are changing has been given better rules for combining independent updates.

That means parallelism is not always a fixed property of the task itself.

Imagine three agents researching vendors before a fourth makes a purchase. If each researcher’s useful progress exists mainly inside a separate browser session, the fourth agent inherits an integration problem. If each researcher instead exports a structured comparison containing everything the purchaser needs, the three browser sessions can disappear. The business question has not changed. The representation of the intermediate work has.

But a portable artifact is not automatically a complete one. A written summary may omit an authenticated session that the next step needs. A file export may omit transactions already made against a remote system. A checkpoint preserves one running process, not the semantic meaning of two processes that changed the same outside world.

The right test is not “was the work saved?” It is “did the saved form preserve everything the next step needs, and does it have an acceptable way to combine with the other saved forms?”

Git provides a useful warning here too. A July preprint analyzed 33,596 agent-authored pull requests, then replayed 747 selected pairs of concurrently active pull requests—one pair per repository. In that replay sample, 41.7 percent of cross-agent pairs produced textual merge conflicts, compared with 19.8 percent of intra-agent pairs. Cross-agent pairs were rare in the full corpus, so those percentages should not be read as the everyday conflict rate for coding agents. They establish a narrower point: an explicit merge operation can still leave expensive reconciliation work.

And even a textually clean merge can be wrong. Software-engineering research distinguishes textual conflicts from semantic ones, where combined code may still fail to compile or pass tests. A study later published at ISSTA 2022 describes semantic merge conflicts in precisely those terms.

Giving work a merge operation moves the boundary. It does not erase the boundary.

You can also move the conflict earlier

Spine-Branch avoids one kind of impossible fan-in by keeping important live state on a single lineage. That is not the only possible design.

A May preprint called STORM attacks a related problem in coding agents from almost the opposite direction. Instead of giving each worker a private Git worktree and postponing conflicts until a final merge, STORM mediates interactions with a versioned shared workspace so conflicting edits can be detected and resolved at write time.

Its authors report an 18.7-point improvement over their Git-worktree multi-agent baseline on Commit0-Lite and a 1.4-point improvement on PaperBench. Those benchmark results are not proof that shared mediation is generally better than branches. They demonstrate a useful rival design: if late reconciliation is costly, a system can sometimes constrain divergence earlier.

Now the problem has three moving parts.

The task determines what can be done independently. The representation determines what can be combined afterward. The coordination protocol determines when conflicts become visible and how much divergent work can accumulate before they must be resolved.

Changing any one of those can change the team that makes sense.

There is still another gate

Suppose the state problem is solved. Every branch returns a clean, composable artifact. Parallel work can still be a bad idea.

It has to save more time or produce more value than it costs to coordinate.

The Spine-Branch authors report that on a separate set of OSWorld 2.0 workflows, multi-agent coordination was often not worth the overhead because many tasks did not contain enough useful independent work. A broader study published in Nature Machine Intelligence in July reached a similar boundary from another direction. Across 260 configurations covering six benchmarks, five coordination architectures, and three model families, single-agent baseline performance was the most robust predictor of whether coordination helped or hurt in the tested domains.

The researchers also fitted a practical threshold rule. It predicted the sign of the multi-agent effect in 15 of 16 validation configurations on SWE-bench Verified and Terminal-Bench. The authors explicitly warn against treating that threshold as a universal scaling law; the underlying baseline-by-team-size interaction did not survive their cluster-robust correction.

So useful parallelism has at least two gates.

First: can the independent work come back together with acceptable fidelity and reconciliation cost?

Second: is there enough valuable independent work for the speedup to exceed the planning, communication, and reconciliation overhead?

A task can pass one gate and fail the other. A perfectly mergeable job may still be quicker for one strong worker. A richly decomposable job may still be a poor candidate for parallelism if the branches leave behind states that cannot be safely recombined.

This is why “add more agents” is such an unstable prescription. Agent count comes after the more basic design decisions.

Look at the arrow coming back in

Draw a workflow as boxes and arrows. Two arrows leaving a box invite an obvious question: can these pieces run at the same time?

The more consequential arrow may be the one where those paths meet again.

Before assigning another worker, ask what each branch will actually leave behind. Is it ordinary information, a versioned artifact, a live environment, a remote side effect, or some mixture? If two branches change it independently, what is supposed to reconcile them? Which conflicts will that operation detect? Which can pass silently? Could important state be exported into a form that rejoins more safely? Would moving the conflict earlier be cheaper? And, after all that, is the parallel work large enough to pay for the coordination?

Most answers will be mundane. Export the file. Commit the patch. Keep one authenticated session on the main path. Reject conflicting writes early. Serialize the payment. Or do not split the work.

A workflow diagram makes the meeting point look like punctuation. It is actually a claim: these two histories can become one without losing what matters.

That claim deserves to be designed before the second branch does any work.

Sources