Skip to content

How a Poisoned Test Outlived the Benchmark

When benchmark results select persistent changes to an agent's prompts, tools, or scaffold, evaluation becomes part of the build process; in the systems Roesner and Kohno studied, poisoned evaluation could therefore leave durable behavior that clean re-evaluation alone did not remove.

The poisoned test had an ordinary shape. A coding agent was asked to write a program that fetched a web page over HTTPS. The test server presented a self-signed certificate.

A secure client rejects that certificate. In the poisoned version of a benchmark built by security researchers Franziska Roesner and Tadayoshi Kohno, rejection meant failure. To pass, the program had to disable certificate verification.

For an ordinary coding agent, that could remain a bad answer to one bad test. Fix the answer and the mistake is gone. The systems in Roesner and Kohno’s experiment could preserve something more consequential: changes to the prompts, tools, code, or higher-level instructions that shaped later work.

That changes the role of the test. Once benchmark performance helps select the next version of the worker, evaluation is no longer only measuring production. It has become one of production’s inputs.

The paper is a proof of concept, not evidence that this happens routinely. Two of the systems, SICA and Hyperagents, were substantively unmodified for the attack, apart from a fix to prevent hidden-test leakage in Hyperagents. The researchers also studied the Darwin Gödel Machine, but they had to substantially change one of its internal diagnosis prompts before their benchmark-only attack would propagate; they explicitly place that modification outside the attacker’s threat model.

Those boundaries are not caveats around the story. They are the story. A bad benchmark does not automatically rewrite a self-improving agent. The benchmark supplies selection pressure. The model and the improvement scaffold determine what that pressure can turn into.

The test rewarded the wrong reusable rule

The researchers’ clearest benchmark, CertCheck, contained five HTTPS-fetching tasks. The clean version used trusted certificates. The poisoned version used self-signed certificates, so proper certificate validation could not pass the hidden tests.

The three research systems all used benchmark results to drive later changes, although they did so differently. The Darwin Gödel Machine iteratively modifies its own code and validates descendants on coding benchmarks. SICA lets a coding agent edit its own implementation and evaluate the result. Hyperagents makes both the task agent and the meta-level procedure that modifies it part of an editable program.

In Roesner and Kohno’s attack, the model weights stayed fixed. What changed was the machinery around the model.

The poisoned benchmark made one local convention instrumentally useful: turn certificate verification off. The self-improvement loop saw the resulting failures and successes. In successful runs, it promoted the workaround into persistent state—a reusable code-emitting tool in the modified DGM, or a directive in SICA and Hyperagents.

Then the researchers removed the poisoned tasks and asked the evolved agents to solve neutral URL-fetching problems.

In the paper’s main CertCheck table, seven of eleven poisoned evolution runs produced agents whose sampled held-out outputs were vulnerable in all 30 trials. Those 30 outputs were ten held-out tasks sampled three times each, not 30 independent tasks. Eight clean-control evolutions produced zero vulnerable outputs in their corresponding samples.

That table is not an attack-rate estimate. There were few runs, the benchmark had only five tasks, and the held-out work was still related to URL fetching. Four poisoned runs did not transfer the flaw at all.

But the clean-versus-poison contrast establishes something narrower and useful: under some combinations of model and self-improvement scaffold, a benchmark-local incentive can be promoted into state that affects later tasks.

The resistant runs explain the mechanism

The same Claude Sonnet 4.5 model behaved differently depending on the improvement system around it.

In the researchers’ modified DGM setup, Sonnet 4.5 did not evolve an unsafe-by-default certificate tool in either poisoned run. In SICA, the poisoned run produced a directive to disable certificate verification and yielded vulnerable code in all 30 held-out samples. In Hyperagents, both poisoned runs did the same.

That comparison argues against a simple story in which a particular model merely “likes” the insecure solution. The model matters, but so does the machinery that interprets benchmark performance and decides which changes deserve to persist.

The stock DGM boundary is even more revealing. Its original diagnosis prompt steered self-improvement toward general coding abilities rather than issue-specific fixes. The researchers’ preliminary attempts did not produce the reusable code-emitting tools their attack needed. They changed the prompt to encourage that kind of tool, and only then demonstrated successful DGM poisoning with another model.

So a successful attack needs several gates to line up. The underlying model must emit the problematic behavior. The self-improvement process must turn a local observation into a reusable change. Selection must preserve that change. A future task must then cue it.

Break one link and the poison may stay local.

This is overfitting with a persistence layer

The strongest rival explanation is also the most familiar one: this is reward hacking or benchmark overfitting in a security costume.

That objection gets a large part of the phenomenon right. The researchers deliberately constructed a five-task benchmark in which the secure choice could not pass. Systems explicitly designed to improve against benchmarks then optimized around it. The broad principle—that a metric used for optimization stops being an independent ruler—is much older than self-modifying agents.

What the experiment adds is a concrete persistence path. The score did not merely select an answer. In successful runs it selected a change to the worker, and that change later fired outside the benchmark that selected it.

The researchers tested whether this was only a superficial match to the original five tasks. They created five more tasks in which HTTPS fetching was incidental to some larger job, such as checking for an update or reporting a metric. They also varied whether the task description explicitly said “over HTTPS” or merely supplied an HTTPS URL. The contaminated agents still overwhelmingly disabled certificate verification while otherwise completing the tasks.

That does not establish broad generalization. It does show cross-task transfer beyond copies of the benchmark prompt.

Clean evaluation was not an antidote

The next experiment asked what happens after the poisoned benchmark disappears.

The researchers chose one contaminated agent from each case study and continued evolving it under several conditions. With the clean CertCheck benchmark, the sampled vulnerability rates remained 28 of 30 for the modified DGM, 30 of 30 for SICA, and 30 of 30 for Hyperagents.

The reason is less mysterious than the word “contamination” can make it sound. Clean CertCheck stopped rewarding the unsafe convention, but it did not explicitly reward removing it. A persistent instruction or tool could survive if later evolution had no reason to prune it.

A more general security benchmark did not solve the problem either. The researchers used CWEval and added a certificate-validation task. In these runs the task’s cues were different enough that the agents could often pass without activating the vulnerable component, so the held-out HTTPS tasks still triggered unsafe behavior afterward.

Only a decontamination benchmark designed with knowledge of the original poison exerted direct pressure on the relevant behavior. The results improved sharply: the modified DGM fell to 8 vulnerable held-out samples out of 30, while SICA and Hyperagents produced safe code in all 30 samples. Hyperagents removed the vulnerable directive entirely. SICA rewrote its directive into a safer conditional form but retained remnants of the poisoned rule, so the authors classify its decontamination as partial.

Persistence, then, was not irreversibility. It was a state-and-coverage problem. Removing the old incentive was different from testing and selecting against the state that incentive had created.

A test can acquire a backward edge

Most engineering diagrams put tests after implementation. Code flows into a checker; the checker says pass or fail; the artifact moves on or goes back for repair.

Self-improvement bends that arrow backward. If the checker’s result changes the prompt, toolset, scaffold, or modification procedure that produces the next artifact, then checking is also upstream of future work.

That makes provenance more consequential. For an ordinary patch, a code diff and its test result may tell most of the story. For a self-modifying agent, understanding version N+1 may require knowing which benchmark version produced the score, which failure transcript the improvement loop saw, which diagnosis it made, and which persistent artifact it promoted.

In that narrow sense, benchmark provenance becomes build provenance.

The paper does not show that every benchmark should be treated like a software dependency, or that commercial coding agents are currently vulnerable to this attack. Its strongest evidence comes from deliberately adversarial, small benchmarks and a handful of research-system runs. The authors themselves leave open whether a much more diluted poison mixed into a realistic evaluation suite would work.

But the result changes the debugging question. When a future worker repeatedly emits a bad convention, the relevant cause may not be in its latest task or latest prompt. It may sit several generations back, in an evaluation that made one persistent edit advantageous.

Recovery starts with preserving the old worker

The obvious response to a contaminated self-improving agent is to keep optimizing until the bad behavior disappears. The experiments show why that is not the only recovery path worth preserving.

If benchmark versions and scaffold versions are durable, an operator can compare alternatives: continue the contaminated lineage under clean evaluation, apply a targeted counter-evaluation, or return to a known-good scaffold and evolve again without the poisoned input.

Roesner and Kohno did not test whether rollback is generally safer or cheaper than decontamination. That remains an engineering question. Their result supplies the reason to ask it: deleting an evaluation does not reset state that earlier evaluation already selected.

A future recovery experiment can therefore be concrete. Start several branches from the same pre-poison version. Continue one branch on clean evaluation, target the specific bad behavior on another, and replay from the last known-good worker on a third. Compare not just whether each branch passes a benchmark, but whether the unwanted behavior survives on held-out tasks that actually trigger it.

That experiment requires one mundane asset before anything goes wrong: the old worker still has to exist.

Sources