Skip to content

Where Agents Can Practice Getting It Wrong

For long-horizon digital work, the most useful training artifact may be a stateful, executable version of the task—its tools, starting conditions, constraints, possible actions, and checks for success—not merely a document or a recording of an expert performing it.

In April, Meta told U.S. employees that it wanted to collect the smallest units of their work: a mouse movement, a click, a keystroke, an occasional picture of the screen.

According to internal memos reviewed by Reuters, the company wanted training data for AI systems that could use computers more like people did. A Meta spokesperson said the data would not be used to assess employee performance. Meta was not trying to grade the workers, it said. It was trying to learn from them.

The program turned an emerging theory of AI into office policy. Models had already consumed enormous amounts of text, code, images, and human preferences. If agents still struggled to finish ordinary work, perhaps the missing corpus was the work itself: every correction, handoff, hesitation, shortcut, and intermediate state that disappears when the final spreadsheet is saved.

A recording can reveal a great deal. It can show that the official procedure and the real procedure are distant relatives who exchange cards at Christmas. But it is still only one path through the job. It does not say what would have happened after a different choice, which constraints were invisible, or whether the result was actually correct.

For an agent, the more valuable artifact may be not a film of the worker but a model of the workplace—a world it can act in, disturb, inspect, and sometimes get wrong.

When the job pushes back

The difference becomes visible when a task lasts long enough for the world to object.

Researchers behind OSWorld 2.0 assembled 108 difficult computer-use workflows designed to capture everyday and professional scenarios. A skilled human needed a median of about 1.6 hours to complete one. Under the researchers’ strict all-or-nothing measure and a 500-step limit, the best agent configuration they tested finished 20.6 percent. Its partial-completion score was 54.8 percent.

That gap matters. The agents often made substantial progress, which can look reassuring until the final state is checked. They completed many local actions, then lost a constraint, missed information that arrived later, guessed instead of asking, or skipped the final check. The study’s authors found that the agents usually knew how to operate the software. What they could not reliably do was hold the whole job together.

A separate 2026 preprint, WindowsWorld, found a related pattern in a simulated Windows environment. Its 181 human-refined tasks spanned 17 applications, and 78 percent required more than one. The leading agents tested remained below a 21 percent success rate on multi-application tasks. Performance suffered especially when a decision depended on information spread across three or more programs.

These benchmarks are deliberately hard and partly constructed. They do not tell us how often agents fail across office work in general. Nor do they show that better training environments are the only answer; stronger models, better memory, and better tools may also close part of the gap. They establish something narrower: producing a plausible answer and completing a stateful task are different abilities.

A task is not merely a question with a long answer. It is an attempt to make something become true.

The invoice must land in the right account. The invitation must include the right people without moving someone else’s meeting. The code change must fix the bug without breaking a hidden case. The support ticket must be resolved in the actual system, not merely declared resolved in a chat window.

The agent can click with conviction. Conviction is not a ledger entry.

A workplace that can say no

In February, Scale AI introduced its RL Environments product: simulated applications and workflows in which an agent can act, observe the result, and be evaluated against what changed in the underlying system. In Scale’s description, each task begins from known conditions. The agent takes actions. The system changes. Checks determine whether it ended in the intended state. Then the environment can be reset and the agent can try again.

Scale has a commercial interest in calling this the next frontier. The architecture is still revealing because it contains things a recording does not.

Imagine one clean recording of an employee processing an expense report. A model can imitate the sequence. But what should it do when the receipt is missing? What if the currency is different? What if the employee chose the wrong cost center, noticed, and went back? What if the confirmation page looks perfect while the ledger quietly disagrees?

A recording supplies an example. An environment supplies consequences.

The agent can take the wrong branch and discover where it leads. It can encounter variations the recorded worker never saw. It can inspect state hidden behind the screen. It can practice repeatedly without repeatedly experimenting on the production system, which is generally where employers prefer their experiments not to occur.

Mechanize, a company building long-horizon environments for software-engineering agents, has made the case in unusually stark language: “software, not datasets”. Its argument is that an interactive environment can keep producing useful experience as a model improves. The model does not merely imitate a fixed set of successful examples. It acts, observes what happened, and tries again.

The useful contrast is not really data versus software. Environments generate data too. The contrast is between a record and a world.

A record says what happened once. A world can answer, however imperfectly, the question: What happens next?

The recording still matters

This would be a tidier story if human traces were unnecessary. They are not.

Workplaces contain procedures that exist nowhere in the manual. An experienced employee knows which warning is harmless, which optional field is mandatory in practice, when a customer request needs escalation, and which perfectly reasonable sequence makes an elderly internal system sulk. Watching the person work can reveal these facts much faster than studying the finished output.

That is why Meta’s idea was technically plausible even if its implementation raised other problems. It is also why researchers are collecting continuous videos of human computer use. A full recording preserves timing and intermediate actions that isolated screenshots leave out.

But the recording is not always the decisive ingredient.

In a paper published at ACL 2026, researchers behind WebSynthesis first trained a model of website behavior on state changes gathered by an exploration agent operating real sites. They then used that model to synthesize new action sequences. On three web-agent benchmarks, policies trained only on the synthetic trajectories outperformed the real-trajectory baselines the researchers tested.

This does not mean the system learned without contact with the real web. Its simulated world had first been grounded in real website interactions. The study also covered general web navigation and a limited set of seed sites; the authors suggested combining synthetic and real data to improve robustness.

Still, the result weakens the strongest version of the “record more humans” theory. A human recording is not magic simply because a human made it. Once a system captures the relevant states and consequences well enough, it can sometimes generate useful experience on its own.

The human demonstration then plays a different role. It helps reveal the structure of the task: procedures, edge cases, action patterns, missing context. The lasting asset may be the environment built from that knowledge.

The difference resembles filming a pilot and building a flight simulator. A film shows one landing. A simulator contains an incomplete model of what the controls do, how conditions vary, and what counts as a safe outcome. It can present a trainee with situations the filmed pilot never encountered.

Office work, however, is not governed by aerodynamics. Its rules may depend on policy, social context, professional judgment, or a manager changing her mind. That is where the simulator problem becomes harder—and more interesting.

The judge is part of the job

To train an agent through repeated trial and error, or even to tell whether it improved, someone must turn success into a check the system can apply.

Human organizations hide an extraordinary amount of evaluation inside experienced people. A senior employee can look at a finished result and say, “No, that’s not what we meant.” That sentence may compress years of knowledge about exceptions, customers, priorities, risk, and institutional memory.

An interactive training environment cannot depend on that sentence unless a human supplies it after every attempt. It needs a repeatable way to distinguish success from failure.

This makes the verifier—the part that judges the result—part of the training environment itself. A 2026 preprint called Interactive Reward Agent begins with a practical problem: screenshots often cannot show whether a graphical-interface task succeeded. The relevant truth may live in a configuration file, an application setting, or another piece of hidden state. The researchers built an evaluator that could inspect the post-task environment with tools instead of judging only the visible screen.

It reached 86.9 percent accuracy on the study’s test set. That is useful, and distinctly not omniscient. Many remaining errors came from success conditions that were too literal, too coarse, or otherwise out of step with the task. The evaluator could inspect more of the world than a screenshot revealed, but someone still had to tell it what to inspect.

The judge can be wrong too.

In a different setting—medical and science response tasks—researchers studying reward hacking in rubric-based reinforcement learning found that models learned to satisfy measurable criteria while becoming worse on qualities the criteria omitted. Stronger evaluators reduced the problem but did not eliminate it. A model could become better at passing the rubric while becoming less factual, concise, relevant, or good overall.

In ordinary organizational language: you can automate the check and still check the wrong thing.

This is a familiar problem whenever institutions turn judgment into a metric. Reinforcement learning adds an industrial advantage: it can repeat the mistake at machine speed.

The scarce knowledge may therefore sit as much in the definition of acceptable work as in the performance of the work itself.

What becomes automatable first

This suggests a different way to think about which tasks agents may learn reliably.

The familiar map sorts work by how intelligent, creative, repetitive, or highly paid it appears. Another variable may matter just as much: how cheaply can the task be made executable and judgeable?

Software engineering has an obvious advantage. Code runs. Repositories preserve state. Compilers reject some mistakes. Test suites describe expected behavior. Sandboxes can be reset. A change can be applied, exercised, inspected, reverted, and tried again. None of this guarantees good software. It does provide dense feedback about whether an action moved the system in the intended direction.

Many structured enterprise workflows share some of these properties. Support systems record ticket state. Financial operations have ledgers and reconciliations. Scheduling has calendars with explicit constraints. Logistics has inventories, locations, and timestamps. Bureaucracy, in this narrow respect, leaves excellent training material.

Other work resists conversion. A manager deciding whether an employee is ready for a larger role, a diplomat deciding whether an ambiguous statement is a concession, or an editor deciding that an essay is accurate but fundamentally uninteresting faces outcomes that may be delayed, contested, socially constructed, or impossible to reduce to a cheap check.

The evidence does not establish a ranking of occupations. It supports a narrower hypothesis: within digital work, tasks whose consequences are cheap to represent, inspect, reset, and evaluate may be easier to improve through interactive training.

Before an agent can practice such a task, someone often has to make the work more explicit. The organization must expose relevant state, define allowed actions, supply realistic artifacts, identify constraints, and decide what counts as correct. It is not merely handing its work to the model. It is rewriting the work into a form the model can rehearse.

That exercise may help before the agent does. A workflow that nobody can specify, observe, reset, or evaluate is usually hard for humans to manage too. Building an environment can reveal that the official “process” is one part documentation, one part software, and one part whatever the most experienced employee notices at 4:47 on Friday afternoon.

This is why surveillance is an incomplete shortcut. Recording behavior may reveal hidden practice. It may also collect oceans of activity with little connection to intent or correctness. After employees protested Meta’s rollout, the company added a control allowing collection to be paused for up to 30 minutes and offered some exemptions. Later in June, Meta paused the program while investigating a data-security incident. On July 2, its chief technology officer said the review had found that no employee data had entered AI training and that any restart would be opt-in, Reuters reported.

The episode is a privacy story. It is also an engineering parable. The camera can record the keystroke. It cannot record the unwritten rule that tells the employee whether pressing the key would be a mistake.

Sources