Clues an Agent Never Sees
An autonomous agent's performance depends on which evidence its trajectory exposes, as well as what the model can infer from that evidence; the ART study demonstrates an observation-acquisition failure mode without isolating it as the sole cause of the direct-context versus tool-use gap.
The first reason an AI worker had found a particular enzyme interesting did not survive investigation.
The worker was surveying reverse transcriptases, enzymes that copy RNA into DNA. A nearby gene looked like a possible partner for one of them. On closer inspection, that association appeared incidental. The worker rejected it but queued another task to examine the reverse transcriptase more closely.
A supervisor then directed the next worker toward the DNA just before the enzyme’s gene, known as its upstream flank. Related enzymes sometimes use an RNA component encoded there. Reading those stretches in the enzyme’s relatives, the worker noticed a repeating pattern.
That sequence of decisions, documented in Anthropic’s September 23 technical report, began the identification of array-associated reverse transcriptases, or ART: a family of phage enzymes associated with an unusual repeat array and a nearby partner gene. Follow-up work found array-derived RNAs in infection data. The researchers still do not know the system’s biological function, whether the enzyme is active, or whether those RNAs are its substrates.
The discovery is unfinished biology. It also left behind a revealing record of how an agent came to notice something.
Anthropic ran the larger search campaign ten more times. Nearly every rerun that completed the initial census encountered ART loci, and two investigated the lineage further. But the authors’ audit found that none read the upstream DNA, and none rediscovered the array. Their search used identifiers from the original campaign, so it could miss loci outside that set.
Reaching the right neighborhood had not reproduced the observation.
The clue behind the tool
A tool-using agent decides what to inspect. A database can contain the right record without the agent querying it. A repository can contain the decisive function without the agent opening its file. Success requires both acquiring the evidence and interpreting what comes back.
The ART researchers tested that distinction with a more focused assignment. Instead of asking models to search the vast biological database, they asked them to characterize the ART system from specified inputs.
Seven Claude models each made 100 attempts at five information levels. The first supplied two enzyme protein sequences. The second placed the corresponding DNA loci directly in the prompt. The third supplied all 96 loci as files in a tool-enabled environment. Later levels added predicted structures, literature search and web access.
For the four strongest models, putting the DNA loci directly in context produced repeat-array recognition in at least 90 percent of attempts. With tools, recognition fell as low as 32 percent for one model and information level.
That outcome was defined by a judge model: it counted whether each final report asserted the authors’ repeat-array claim. It was not a blinded human assessment or evidence that the model had explained ART’s function.
The researchers also inspected what the agents actually read. Across the four strongest models’ file-based attempts, 39 percent never brought a contiguous stretch of at least 200 nucleotides into context. That meant seeing no more than about one repeat unit. For each of those models, attempts that read at least that much sequence recognized the array 16 to 32 percentage points more often.
The distinction was visible in the trajectories. A model that could describe a repeated pattern when shown the DNA often did not fetch enough of it when given responsibility for the investigation.
But the experiment does not isolate a single cause. Moving from two loci in a prompt to 96 loci in files changes the amount of information, navigation burden, ordering and available actions. Reading more DNA is associated with recognition; it was not randomly imposed on otherwise identical runs. The same-family judge adds another measurement limit.
What survives those qualifications is a concrete failure mode. Some attempts reached their final report without exposing much of the raw sequence that the direct-context condition had made easy to inspect.
What access leaves undecided
Tools are often described as giving an agent more context. More precisely, they give it more possible actions.
Between a useful file existing and its contents entering a model’s view, the agent must decide where to look, which operation to use and how much to retrieve. A search result might identify the right object without revealing its important feature. A short preview might show an isolated repeat without enough surrounding sequence to make repetition apparent.
The successful ART run depended on a particular chain of attention. A rejected partner-gene hypothesis did not end the investigation. A supervisor’s follow-up pointed toward upstream RNA. The next worker opened the raw DNA of related loci. Only then could the repeated structure become evidence.
This suggests four distinguishable questions: Was the evidence accessible? Was it selected for inspection? Did enough of it enter context in a useful form? Was it interpreted correctly?
Those are analytical distinctions, not four independently measured components of intelligence. An agent can move back and forth among them. Interpreting a partial result may determine the next search. Nevertheless, the distinctions matter when deciding what to improve.
If a model reads the relevant sequence and misses the repeats, better interpretation or a clearer representation may help. If it never reads the sequence, improving recognition alone leaves the acquisition problem in place. A different search policy, more informative preview or targeted instruction might help, but the ART study does not establish which intervention would work best.
Nor does it show that a richer environment is generally harmful. In the same benchmark, additional tools and information helped the stronger models describe the partner gene. The environment could help one part of the investigation while leaving another unseen.
Finding the object is work
Software agents encounter a related problem in a more familiar setting.
The SWE-agent project treated the interface between a language model and a codebase as something to design and test. Its file viewer limited what was shown at once; its search tools controlled how results reached the model. Changing those interfaces changed issue-resolution performance without changing the model’s weights.
The result did not reduce to showing more code. In its reported viewer ablation, a full-file view performed worse than a 100-line window. The relevant design problem was what the agent could use effectively while navigating, not simply how much text the system could display.
More recent work makes finding code a separate assignment. CodeScout trains an agent to locate the files, classes and functions relevant to an issue. Its researchers also tested supplying those locations to agents attempting the repair. That makes evidence acquisition an explicit part of the workflow rather than assuming it will happen as a side effect of solving the problem.
These are independent software experiments, not replications of ART. They support the narrower proposition that an interface and a search process help determine which evidence a capable model reaches.
The underlying problem predates language models. Research on information foraging studies how people follow imperfect cues toward valuable information while paying the cost of navigation. Agents make this familiar work unusually easy to hide: an environment may expose hundreds of tools and millions of records, while a performance summary shows only the final answer.
Two agents can therefore receive the same failing score for different reasons. One inspected the faulty function and formed the wrong explanation. The other never opened it. The score records task failure correctly. It does not tell an engineer which problem to work on.
Once the evidence arrives
Exposure still does not guarantee successful use.
In Lost in the Middle, models answered differently depending on where relevant information appeared within a long context. Material can already be in the prompt and remain underused. An observation trace can establish that text was presented; it cannot by itself establish that the model understood it.
This is why the ART benchmark should not become a rule to dump more raw data into every agent’s prompt. The targeted sequence was especially useful because the task concerned that biological system. In a broad campaign, most available sequences will not contain the sought-after clue. Inspection costs time and can crowd out other useful work.
The useful question is where a particular trajectory lost the opportunity to succeed. Was the source unreachable? Did the agent choose not to inspect it? Did the tool return too little? Or did interpretation fail after adequate exposure?
Answering those questions requires a record of observations, not just a final report: which objects were opened, which regions were returned and what the model had received before reaching a conclusion. Even then, traces support diagnosis rather than automatically proving causation. The ART study’s reading threshold remains an association, and its direct-versus-files comparison changes several conditions at once.
The successful campaign gives that limitation a concrete counterpart. The initial partner-gene theory was discarded. A follow-up nevertheless brought an upstream stretch of DNA into view, where its repeated structure could be recognized.
The reruns could reach the same lineage without making that observation. What needed repeating was not merely the search for the enzyme. It was the decision to look beside it.
Sources
- Peter H. Yoon, Januka S. Athukoralage, Emmanuel Ameisen, Eric Kauderer-Abrams, Nicholas T. Perry, and Matthew G. Durrant, Autonomous AI Agents Discover Reverse Transcriptases with Tandem Repeat Arrays, September 23, 2026.
- John Yang et al., SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, 2024.
- Lintang Sutawika et al., CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents, March 18, 2026.
- Amy J. Ko, Brad A. Myers, Michael J. Coblenz, and Htet Htet Aung, An Exploratory Study of How Developers Seek, Relate, and Collect Relevant Information during Software Maintenance Tasks, IEEE Transactions on Software Engineering, 2006.
- Nelson F. Liu et al., Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2024.