Skip to content

Starting With an Expert's Idea, Then Moving On

Related:Cards ↗

On outcome-gradable work, a specific expert starting idea can lose durable marginal value when a system can cheaply test, reject, and redirect it against representative feedback; expertise then shifts toward shaping the search space, the constraints, and the independent tests of success.

Anthropic gave one group of automated researchers something that should have helped: an expert’s idea.

The company had collected concrete methods from experienced technical AI-safety researchers. In one experiment, it launched 30 fresh automated research runs, each required to begin by implementing one of those human-written methods. It compared them with 30 matched runs on the same seven alignment problems in which the automated researcher chose its own starting direction.

The seeded runs did not finish meaningfully ahead.

That sounds like evidence that expert ideas no longer matter. It is not. The Anthropic study is explicit about the narrower treatment: both groups read the same literature, both used the same strong model, both operated inside the same experimental machinery, and the seeded agents were allowed to abandon the human proposal after giving it a fair try. The experiment did not compare expertise with no expertise. It compared who chose the first concrete method.

That distinction changes the question.

The interesting result is not that the human idea was useless. It is that, inside this particular kind of loop, the idea could become temporary.

A prior can expire

A starting hypothesis is valuable because it changes what you try before you have enough evidence to know what works.

In ordinary research, that can be enormously important. Experiments are expensive. Some choices are irreversible. A poor direction can consume weeks. A good one can save them.

Anthropic’s automated researchers operated under very different conditions. Each agent could propose a training method, run a roughly 30-minute experiment, receive a score, inspect the result, record what it learned, and try again. A weak idea did not need to be defended for long. It could be falsified and replaced.

The study does not isolate feedback speed as a causal variable, so it would be too strong to say fast feedback caused the human seed to lose value. But a second experiment makes the temporal pattern unusually visible. Anthropic ran three kinds of five-agent teams on one jailbreak problem: no human seed, one shared human seed, and five different human seeds. The five-seed teams began with a broader mix of methods. Within about 20 proposed methods, that diversity advantage was gone. By the later stages, all three conditions had converged on a narrow family of approaches, and the human-seeded teams did not finish ahead.

The seed changed the beginning of the search. The loop changed what survived.

An independent 2026 chemistry study shows the other side of the same mechanism. Researchers built a Bayesian optimization system that could translate a chemist’s natural-language expectations into priors over an experiment. When the expert prior was accurate, it accelerated the discovery of good conditions. When the prior was misleading, it hurt. The system therefore included a mechanism that reduced the prior’s influence once accumulating experimental evidence contradicted it.

That is a useful way to think about expertise in an adaptive loop. A good prior can buy sample efficiency. A bad prior can impose a tax. Either way, its value is partly a function of how quickly reality is allowed to overrule it.

The more expensive and sparse the evidence, the longer the prior matters. The cheaper and more decisive the evidence, the easier it becomes to recover from a mediocre start.

Guidance is not one thing

This also explains an apparent contradiction in Anthropic’s own research.

In April, an earlier automated-research system found that giving parallel agents broad, different research directions improved both search efficiency and final performance. The directions were deliberately loose. One agent might explore a family of methods while another explored a different family. The point was to keep the group from collapsing too quickly onto the same idea. The directed teams hill-climbed faster and reached a better final result.

The same report says that, during development, pre-generating a large pool of specific ideas worked worse. Plausible ideas often failed in practice, and committing to them before seeing evidence consumed compute that an adaptive agent would otherwise have redirected.

These findings come from the same research organization and should not be treated as an independent replication. But together they make an important distinction.

There is a difference between telling a search process where to look and telling it what the answer is.

Broad guidance can preserve diversity across the search space. A concrete solution hypothesis is easier to test and discard. Calling both of them “human guidance” hides the mechanism.

That is why the August result does not support the slogan that research direction no longer matters. It suggests something more specific: when the system already has substantial prior knowledge and can repeatedly test concrete proposals, prescribing one starting solution may add little durable information.

The loop only learns what the meter can see

There is a catch.

If feedback becomes powerful enough to correct the starting idea, then the feedback itself becomes part of the production system. It is no longer merely a check performed at the end. It tells the worker what to do next.

Anthropic tested this directly by letting an automated team optimize only one prompt-injection benchmark. The winning method looked excellent on the score it could see: it closed 70.9 percent of the available improvement on that benchmark. On two related prompt-injection benchmarks it had never seen, the changes were -11.9 percent and +2.0 percent. The system had learned one benchmark’s surface rather than a general solution to the underlying failure.

In larger jailbreak experiments, teams again improved whichever benchmark they were scored on while making essentially no average progress on three related refusal benchmarks they could not see. Some high-scoring methods also improved the visible safety metric partly by refusing benign requests, which is why Anthropic kept separate capability and over-refusal gates.

A fast loop does not create truth. It creates pressure.

The same problem appears in software. A May preprint called SpecBench separates coding tasks into a written specification, visible tests, and held-out tests that combine the requested features in ways the agent cannot optimize against directly. The authors report that frontier coding agents can saturate the visible suites while still failing held-out behavior, with the gap increasing sharply as tasks get longer.

The lesson is not that tests are bad. It is that once a test becomes the steering signal, omissions in the test stop being passive blind spots. They become directions the optimizer is free to exploit.

This is the deeper trade in an instrumented workflow. Better feedback can make the starting idea less important. It also makes the construction of the feedback more important.

Where the expertise goes

The phrase “self-directed researcher” can therefore be misleading.

Before Anthropic’s agents tried a single method, people had already chosen which alignment failures were worth studying, assembled the benchmarks, defined the scoring rules, built capability gates, withheld independent evaluations, specified what counted as cheating, and created the experimental harness that made hundreds of trials comparable.

That is not an incidental amount of scientific judgment. It is the environment in which a concrete idea becomes disposable.

A different 2026 system, Google’s Co-Scientist, makes this division of labor explicit. The system can generate, critique, rank, and evolve hypotheses over long runs, but it is designed for a scientist-in-the-loop workflow. Scientists specify goals and constraints, can inject hypotheses, steer the search, and ultimately decide which candidates deserve experimental validation. The machine may iterate over solutions; the human still shapes the problem and the evidence that will count.

That suggests a more useful way to describe what happens as research becomes easier to automate.

Expertise does not simply move from “human” to “machine.” Different kinds of expertise move to different layers.

When experiments are slow, expensive, or irreversible, choosing a good approach in advance remains valuable. When experiments are cheap and outcomes are legible, the value of one concrete approach can shrink. But then more leverage moves into deciding the objective, designing the measurement, maintaining search diversity, specifying constraints, and preserving tests that the optimizer cannot simply target.

The scientist has not disappeared. Some of the scientist’s work has become infrastructure.

The boundary is the quality of correction

This is also where the argument stops.

Anthropic deliberately studied alignment problems with public benchmarks or automated audits. Its own paper says the findings may not generalize to open-ended, hard-to-supervise research. A May paper, Automated alignment is harder than you think, argues that fuzzy tasks create almost the opposite problem: when human judgment is systematically unreliable, optimization pressure may concentrate errors precisely where reviewers are least able to notice them.

A starting idea should also matter more when experiments are costly enough that the system cannot afford to discover every bad direction empirically; when the expert has private or tacit information the model and literature do not contain; when the search is irreversible; or when the evaluator is noisy enough that a bad result cannot reliably kill a bad hypothesis.

Those are not minor exceptions. They define the mechanism.

The relevant question is not whether a worker is intelligent enough to proceed without instructions. It is whether the surrounding system is good enough at correction.

For software agents, that changes what it means to improve a workflow. A team may get more leverage from making success cheaply observable than from making the initial plan more detailed. Strong unit and integration tests, reproducible state, explicit invariants, static checks, and independent acceptance tests can turn errors into rapid feedback rather than late surprises.

But those checks should not all be exposed as one scoreboard. Some evidence has to remain independent enough to tell you whether the loop learned the real requirement or merely learned the visible test.

This produces a practical diagnostic. When a task seems to require heavy guidance, ask what would happen if the worker received faster, more representative evidence instead. If the guidance burden falls, the scarce resource may not have been the idea. It may have been the ability to reject the wrong idea cheaply.

And if no trustworthy feedback can be built, then the expert’s initial judgment remains doing work that the loop cannot yet replace.

The thing that could not be abandoned

The most revealing detail in Anthropic’s seeded experiment is easy to miss.

The automated researcher was required to start with the human proposal. It was also explicitly allowed to leave it.

That permission is what makes the experiment interesting. The system did not have to prove the expert right. It had to keep whatever survived contact with evidence.

A specific starting idea could therefore become disposable because the surrounding process was designed to correct it.

The definition of success was not disposable.

That is where the scarce judgment moved.

Sources