Skip to content

Building a New Workspace From an Old Log

Incomplete agent traces can constrain the synthesis of useful executable substitutes, but their value is regeneration rather than faithful restoration, and current evidence does not isolate how much of that value comes from the trace rather than the generation pipeline around it.

The first workspace is almost empty.

Researchers take an old terminal-agent trajectory—the sequence of file reads, edits, commands, and observations left behind by a completed run—and replay what it reveals about the state that existed before the agent changed it. In the terminal portion of their experiment, the resulting workspace contains an average of 2.9 files.

Then another model goes to work. It creates missing files, fills partial ones, restores dependencies, and tries to supply the project context the trace never exposed.

The average grows to 22.4 files.

That gap is the central fact in Terminal-Universe, an arXiv preprint submitted September 3 by researchers affiliated with Alibaba’s Qwen team and Tsinghua University. The paper is described as a system for turning agent trajectories into terminal environments. But the numbers make clear that “turning back” is not quite what happens.

Most of the reconstructed world was never in the log.

The missing nineteen files

A work trace can reveal surprisingly concrete facts about a vanished workspace. If an agent reads a file before patching it, the old text may survive in the trace. A command exposes expected paths and tools. An import names a dependency. A failing test reveals something about the surrounding program even when the surrounding program was never opened.

Every action rules out some possible prior worlds.

It does not identify one world uniquely.

The authors say this directly. Their recovered workspace only approximates the original because unaccessed files, implicit system dependencies, and external network resources leave no direct trace. The system therefore has two distinct jobs. First it replays the state that can actually be inferred from the trajectory. Then a stronger model performs what the paper calls agentic completion: it invents enough missing context to make the recovered task workable.

The second step does most of the visible work.

Among the paper’s evaluated terminal workspaces, deterministic replay alone was judged sufficient for the recovered task 40.2 percent of the time. After model-generated completion, the rate rose to 93.5 percent. For the software-engineering pool, the corresponding figures were 20.1 percent and 77.1 percent.

The word sufficient matters more than the word recovered. The judge asks whether a capable agent has enough source, configuration, data, and structure to attempt a particular task. It does not ask whether hidden files match the vanished original. A workspace can therefore pass while being historically wrong.

The paper retains 37,273 task-sufficient environments after its reconstruction, filtering, decontamination, and deduplication pipeline. In a manual check of 30 completed terminal environments, 22 added only task-relevant supporting files; eight added substantial unnecessary files or code. There is no whole-project fidelity measurement against the lost originals.

So a trace is not a compressed backup. It is closer to a set of constraints from which another executable world can be synthesized.

That sounds like a downgrade until you ask what the new world is for.

The old solution can be worse than the old evidence

An agent trajectory contains at least two kinds of information. It records behavior: what one particular policy tried. And it records state evidence: clues about the environment in which those attempts occurred.

Terminal-Universe contains an unusually revealing test of the difference.

The researchers fine-tuned Qwen3.5-27B directly on 35,800 source trajectories. Across two Terminal-Bench 2.1 evaluation scaffolds, the untreated base model averaged 47.0 percent. Direct imitation of the source trajectories lowered the average to 36.7 percent.

The researchers then kept the historical tasks but changed what they learned from them. They reconstructed executable workspaces and had the stronger Qwen3.7-Max teacher solve the tasks again. Fine-tuning on those new solutions raised the average to 52.1 percent.

That result is easy to overread. It does not show that old trajectories are generally bad training data. It shows something narrower and more useful: in this setup, the recorded behavior was a poor thing to imitate, while the recorded interaction still helped bootstrap a process that produced better supervision.

A second ablation points in the same direction. With the recovered tasks and teacher configuration held fixed, training on replay-only environments scored 48.7 percent under the Terminus2-XML scaffold. Giving the model-completion step a chance to fill the missing environment raised the score to 52.9 percent.

The broader pipeline also produced useful downstream data. Under the paper’s full mixture, Qwen3.5-27B improved by 11.9 points on Terminal-Bench 2.1 under Terminus2-XML and by 10.4 points under the Claude Code scaffold; the paper reports a 13.8-point gain on EvoCode-Bench v2 MT@4.

Those gains validate the usefulness of the pipeline. They do not validate the historical accuracy of the generated workspaces.

And they do not tell us how much credit belongs to the trace.

A system built to remember behaves differently

Software has long known how to make execution history useful for reproduction. The difference is that those systems decide what must be remembered before the original environment disappears.

In 2013, the researchers behind ReproZip described a system that traces operating-system calls to capture data dependencies, libraries, configuration, binaries, and other state required to package a computational experiment for execution elsewhere. Record-and-replay systems make a similar bargain: observe the nondeterministic inputs that matter, then preserve them deliberately enough to reproduce an execution later.

An ordinary coding-agent trajectory was not designed for that job. It records whatever the agent happened to expose while solving something else. The distinction is the difference between an authoritative persistence layer and an observation layer.

That is why Terminal-Universe is more interesting as regeneration than restoration. The system accepts that the observation is incomplete, generates a plausible surrounding context, and then asks whether the result can support useful new work.

This also explains where the approach should fail. A transcript may never reveal a database snapshot, a secret, a running service, an external side effect, an unobserved generated artifact, or the exact system package that made a command succeed. If historical identity matters, those things have to be captured by some other mechanism. No amount of eloquent provenance can infer a credential that was intentionally never logged.

The paper itself is careful about a different boundary. Its source traces are selected: the observed end state must expose at least five files and 100 lines before reconstruction begins. The resulting terminal pool is heavily Python-dominated. This is evidence about trace-rich coding and terminal sessions, not arbitrary work logs.

There is another source of uncertainty inside the pipeline. Qwen3.7-Max drives the model-generated stages, including generation and verification. The authors note that a shared teacher can create correlated blind spots: a faulty generated task or solution may also receive tests that fail to expose the fault. The preprint is days old and has not been independently replicated.

The rival does not need the log

If executable environments are the valuable object, perhaps history is unnecessary.

Several neighboring systems start elsewhere. TerminalTraj begins from repositories, builds 32,000 Docker images, and produces 50,733 verified trajectories. CLI-Universe starts from a capability taxonomy and realizes generated tasks in verified Docker environments. CLI-Gym begins with healthy environments and deliberately drives them into failing states to create repair tasks.

They all demonstrate the same broader point: an executable world can be more reusable than one frozen demonstration because it can generate fresh actions and fresh feedback.

They also expose the unresolved question at the center of Terminal-Universe. Does seeding that world from a real trajectory make it cheaper, more realistic, or more useful than generating a world by another route?

The paper does not provide the decisive comparison. There is no matched condition with the same teacher, compute budget, number of environments, and task complexity in which one set of worlds is seeded from trajectories and another is synthesized without them.

That leaves a strong rival explanation for the downstream gains. Perhaps the important ingredients are a powerful teacher, executable filtering, and a large supply of diverse environments, while the trajectory is mainly a convenient seed. The existing ablations show that completion matters once a trace-derived task has been chosen. They do not isolate the trace’s marginal value against an equally resourced alternative source of worlds.

This uncertainty changes the practical lesson. It is premature to declare work logs a new form of capital merely because a sophisticated pipeline can extract value from them. The relevant question is what the historical evidence contributes that cheaper or more direct state capture and environment synthesis cannot.

What to preserve depends on what you want back

A work record can serve three very different goals.

For audit, you want evidence of what happened: who or what acted, what it saw, what it changed, and why a reviewer accepted the result.

For recovery, you want enough authoritative state to resume or reproduce the same work. Git objects, database snapshots, environment manifests, dependency locks, and captured external state matter more here than a readable narrative does.

For regeneration, exact identity may be unnecessary. You need enough constraints to build another executable setting in which related behavior can be tested again.

Those goals reward different retention policies. Saving every virtual machine would preserve far more than most teams need and cost far more than most teams want to pay. Saving only prose can omit the state that later matters. Version control preserves some classes of state with extraordinary fidelity while ignoring others completely. Tool traces preserve a different slice.

Terminal-Universe does not identify the cheapest boundary. It shows why the boundary is worth naming.

After deterministic replay, the average terminal workspace in the paper contains 2.9 files. After a model fills the gaps, it contains 22.4. If you need yesterday’s exact world, that difference is a warning.

If you need a new world that can answer another question, it may be the opportunity.

Sources