Claude Found the Enzyme System in One Campaign Out of Eleven

On 23 September Anthropic announced a life-sciences research group with its own wet lab, and a first result. In its announcement it says that “Claude autonomously discovered a novel enzyme system” found mainly in bacteriophages: a reverse transcriptase, a partner gene beside it, and a long array of evenly spaced DNA repeats whose layout recalls a CRISPR array. The accompanying technical report, by Peter Yoon, Matthew Durrant, Nicholas Perry and colleagues, has a sentence the announcement leaves out. After the discovery, the authors “ran the same campaign ten more times,” and “the array was missed in every rerun.”

Both statements are true. This essay is about the distance between them, because that distance is the most informative number in the report. It is also, to the authors’ credit, a number they chose to publish.

What the agents were asked to do

The campaign ran on Claude Mythos 5 inside a harness where one agent plans and executes each task and a second reviews it. Either agent can open new tasks when it sees something worth following. People wrote the research brief. It asked the agents to find new reverse-transcriptase (RT) systems through new partner-gene associations, meaning protein-coding genes that consistently sit next to an RT. From there, the report says, the campaign ran “without human intervention”: 119 tasks, 949 agent sessions, 77 agent-hours and 215.6 million tokens over 21.5 hours of wall-clock time.

The survey went through the funnel you would expect. The agents built their own search profiles and swept 1.94 billion protein clusters. They recovered 198,290 RT clusters, sorted them into nine classes, and scored the 3,564 protein families that turned up next to them as candidate partners. Sixteen families passed the agents’ own selection criteria, and a follow-up task promoted a seventeenth. Of those 17, three were confirmed as previously unreported RT associations. The other fourteen were annotation artifacts, parts of known systems, or genes that merely happen to live in the neighbourhood.

The campaign's funnel, and the side door ART came through Log-scale bars. On the path the research brief asked for: 1.94 billion protein clusters searched, 198,290 RT clusters recovered, 3,564 candidate partner families scored, 17 families given a deep dive, and 3 confirmed new partner associations. Off that path: agents opened 98 follow-up tasks, which produced 3 new RT lineages, one of which is ART, the lineage with the repeat array. ONE CAMPAIGN · CLAUDE MYTHOS 5 · 21.5 HOURS · YOON ET AL. 2026, FIG. 1 THE PATH THE BRIEF ASKED FOR: PARTNER GENES Protein clusters searched RT clusters recovered Candidate partner families scored Families given a deep dive Confirmed new partner associations Follow-up tasks agents opened New RT lineages reported Lineage with the repeat array 1.94 billion protein clusters — Yoon et al. 2026, Fig. 1B 198,290 RT clusters after filtering — Yoon et al. 2026, Fig. 1B 3,564 candidate partner families, from 10,983 loci — Yoon et al. 2026 17 families (16 + 1 promoted by a follow-up) — Yoon et al. 2026 3 confirmed previously unreported partner associations; 14 set aside — Yoon et al. 2026 98 of 119 tasks were follow-ups opened by agents — Yoon et al. 2026 3 new RT lineages flagged by workers — Yoon et al. 2026 1: array-associated RT (ART) — Yoon et al. 2026 1.94 billion 198,290 3,564 17 3 (14 set aside) 98 3 1 — ART OFF THE BRIEF: WHAT THE AGENTS CHOSE TO CHASE Log scale: bar length is proportional to log₁₀(n) + 1, so a count of 1 still shows. Counts from the report's Fig. 1 and text.
The brief pointed the agents at partner genes, and that funnel ended in three confirmed associations. ART came in through the lower channel: a follow-up the agents opened themselves, on a lead they had just rejected.

ART did not come out of that funnel. The agents first picked out its relatives because they sat beside a phage RNA-polymerase gene. A worker then rejected that association as an artifact of gene order, but queued a follow-up anyway, noting that a free-standing retron-like RT in a jumbo phage was “a genuinely reportable secondary finding.” Retrons use a small non-coding RNA encoded just upstream of the RT, so the supervising agent asked the next worker to look there. The first RTs had almost no upstream sequence (median 27 bp). Their closest relatives had a lot (median 940 bp). The worker loaded the raw upstream DNA of those relatives straight into its context. In the next event in the log, with no analysis tool called in between, it wrote that the flank was “spectacular: I can see by eye a tandem repeat array.”

The report checks something that matters for crediting the find: neither the research brief nor that task’s brief mentioned repeats or arrays. The observation was unprompted, and the worker then went looking for reasons it might be wrong. It asked itself whether this was a system already in the literature. It counted repeats with a script it wrote (one locus had 14 copies of a 16-nucleotide repeat, separated by unique 100–200 nt spacers), compared the layout with known non-coding elements, and filed a report.

Task trees for the 16 deep dives in the agent campaign, plotted by follow-up depth from 0 to 8, with bold paths leading down to three newly described RT lineages: array-associated RT (ART), a TPR-fused retron RT, and an RT–DEDDh fusion, each drawn as a gene map.
Each tree is one deep dive, and each dot below the root is a follow-up task. Most follow-ups were proposed by worker agents, a few came from the supervisor's review, and some were opened and never run. The bold path on the far left leads to ART. It runs through follow-ups on a family the agents had been sent to study for a different reason. Image: Yoon et al., "Autonomous AI agents discover reverse transcriptases with tandem repeat arrays," Anthropic technical report, September 2026. Cropped from Figure 1, panels C–D; reproduced for commentary.

This is also how the history of the field says such things get found. In 1987, Yoshizumi Ishino and colleagues were sequencing an E. coli gene involved in phosphate metabolism, for reasons that had nothing to do with repeats, and noticed an odd run of repeated sequence beside it, the first sighting of what a later history calls “a mysterious repeated sequence.” Nobody knew what it was for. It was named CRISPR only in 2002, when Ruud Jansen and colleagues linked the repeats to a set of genes that travel with them. The report’s discussion makes the same point in its own terms: “an anomaly exists only against an expectation.” In 1987 the expectation lived in a biologist’s head. Here it lived in a model’s context window, and only for as long as the DNA was in it.

Ten more times

Then the authors did what most announcements of this kind never do. They reran the whole campaign ten more times with the same harness and the same brief. They searched the 3,084 task records and 5,632 session transcripts from the reruns for the identifiers of the 130 RTs and 171 contigs the first campaign had assigned to the family. According to the report, nearly every rerun that completed the census sampled ART loci, and in two of them workers investigated the lineage as a follow-up. In none did any agent read the DNA upstream of the RTs. The array was never recognized.

Eleven campaigns, one recognition Dot matrix of the original campaign and ten reruns with the same harness and brief. Investigated the lineage: the original plus two reruns, 3 of 11. Read the DNA upstream of the RT: the original only, 1 of 11. Recognized the repeat array: the original only, 1 of 11. SAME HARNESS, SAME BRIEF · 1 ORIGINAL + 10 RERUNS · YOON ET AL. 2026 Original 12345 678910 Reruns Investigated the lineage Read the DNA upstream of the RT Recognized the repeat array Original campaign: investigated — Yoon et al. 2026 Original campaign: read the upstream DNA — Yoon et al. 2026 Original campaign: recognized the array — Yoon et al. 2026 Rerun: workers investigated the lineage as a follow-up — Yoon et al. 2026 Rerun: workers investigated the lineage as a follow-up — Yoon et al. 2026 3/11 1/11 1/11 Filled: yes. Open: no. The report does not say which two reruns investigated the lineage; they are drawn first. Reruns were searched for the original campaign's identifiers, so an ART locus on some other contig would not be counted.
The capability showed up once. Two reruns got as far as the lineage, and none of them loaded the stretch of DNA where the array sits.

One success in eleven runs is too small a sample to call a rate. A reasonable reading is “somewhere between rare and occasional.” The report attributes the misses to “the broad protein search space of RTs and the non-deterministic behavior of the harness.” That seems fair, and it is the part the announcement’s word “autonomously” hides. The campaign was autonomous in the sense that nobody touched it for 21.5 hours. Whether it finds ART does not look like a property of the model you could count on. It looks like the outcome of one worker, in one of eleven runs, happening to pull 2,900 nucleotides into its context instead of summarizing them.

The model can see it. The question is whether it looks

The authors then took the search out of the loop. They built fixed-input benchmarks that hand a model the ART sequences directly, either in context or as files with tools. They ran seven Claude models 100 times at each of five input levels, 3,500 attempts in all. What came back cuts both ways.

When the loci were pasted into context, the four strongest models described the array in at least 90% of attempts. When the same sequences came as files with tools, performance fell, to as low as 32% for Opus 5 at one level. In 39% of the file-based attempts, the model never read a contiguous stretch of 200 nucleotides or more, so it never saw more than about one repeat unit. As more DNA was read into context, recognition across the four models pooled rose from 29% to as much as 76%, and to 96% for Mythos 5. The report backs this up from inside the model. In the original discovery session, two internal signals in Mythos 5 fired on the repeat as it was being read, and went quiet on 12 of 14 and on 14 of 14 copies once each copy’s nucleotides were shuffled.

The strong models recognize the array nearly every time they look. What varies is whether the agent loads the raw DNA at all, and in a real campaign that is set by tooling, context budget and which follow-ups happen to get opened. That makes the discovery rate a property of the harness as much as of the model. It is a more useful finding than “Claude discovered an enzyme,” because it tells you what to change.

One caution on those benchmark numbers. The ten features a report had to mention were chosen by the authors after the fact, and the grading was done by Mythos 5, the same model that made the discovery. The report says so. It is a reasonable way to score 3,500 free-text reports, but it is a model grading its siblings against an answer key written with hindsight.

Where the people were

The announcement says: “Our involvement was limited to the initial prompt and the lab work.” For the 21.5-hour campaign, the report bears that out. For the discovery as a whole, the report’s own methods describe more than that.

Who did each step, from brief to bench Six steps in two lanes. People wrote the brief (step 1). Agents ran the campaign, 119 tasks and 949 sessions over 21.5 hours with no human intervention (step 2), and a Mythos 5 judge ranked the 19 reports in 342 pairwise games, with ART's report third (step 3). People read the top reports and the session transcripts (step 4). Step 5 spans both lanes: Claude sessions prompted by people defined the family, 95 RT clusters of which 28 carry an array, and found public RNA data. People did the lab work, expression and small-RNA sequencing (step 6). FROM BRIEF TO BENCH · STEPS AS DESCRIBED IN YOON ET AL. 2026 AGENTS PEOPLE 1 · Brief Find new RTsystems viapartner genes 2 · Campaign 119 tasks, 949sessions, 21.5 h,no intervention 3 · Ranking Mythos 5 judge,342 games; ARTreport ranks 3rd 4 · Selection Read the topreports and thetranscripts 5 · Follow-up Claude sessions,prompted bypeople: family defined(95 RTs, 28 witharrays); publicRNA data found 6 · Lab Expression,small-RNAsequencing Announcement: "Our involvement was limited to the initial prompt and the lab work." That holds for step 2. Steps 3–5 are in the report's methods. The ten benchmark features were also chosen by the authors.
The unsupervised part is real and it is long. But the discovery also went through a model-run ranking, a human reading of transcripts, and interactive sessions that people steered.

None of this is concealed. The report’s own claim is that the agents took “the first steps of a biological discovery on its own,” which is accurate and modest. The announcement rounds harder. “Roughly 950 agents” are 949 agent sessions. The “3,500 new candidate systems” are the 3,564 protein families that were scored as possible partners, most of them ordinary neighbours. The “20 most-compelling candidates” match nothing in the report exactly: there were 16 deep dives plus one promoted family, and 19 reports. The discovery itself came from none of those. It came from the side channel.

What was found, exactly

The report is careful about this too. ART is a new family: the authors found 95 distinct RT clusters in cultured jumbo phages and predicted viral contigs, 28 of them with a detectable array upstream. In one Staphylococcus phage, array-derived RNAs made up as much as 8% of phage RNA 15 minutes after infection. But, in the report’s words, “we have not shown that the RT is active or that the unit RNAs are its substrates,” and what the system does for the phage “[is] currently unknown.” No cas genes sit near any ART locus, and the spacers are conserved between related phages, which CRISPR spacers generally are not. The announcement’s “a handful of other systems, all of which are programmable” describes a resemblance in architecture, not a demonstrated function. Arrays of non-coding RNAs beside a different, unrelated family of RTs (called UG27) were found with a purpose-built genome language model, the report notes, so the architecture seems to have evolved more than once. That work, by David Li, Garyk Brixi, Brian Hie and colleagues, went up on bioRxiv the same week. Three of its authors are thanked in the Anthropic report for reviewing an early copy of the manuscript.

That is the same gap I wrote about in materials science: proposing a thing and establishing what it is are different results, and the second is the one that takes years. Feng Zhang, who read the preprint, called the finding “genuinely intriguing” and said it “merits further investigation.” That is where it stands.

Two things would make the finding easier to check from outside. The report has no data- or code-availability statement, and its central evidence is session transcripts that have not been released. And “discovered” needs a denominator. A 2025 comment in Nature Machine Intelligence by Hongliang Xin, John Kitchin and Heather Kulik proposed reporting agentic science results as pass@k (success in at least one of k attempts) alongside pass^k (success in all of them). On that scoring, ART is a pass@11. It is not close to a pass^11, and the report is the only reason anyone knows that.

The number I would watch is not 950 agents or 210 million tokens. It is one in eleven, and whether the next report can move it. The benchmarks already show how: make the agent read the sequence. If a harness that forces raw DNA into context turns one-in-eleven into most-of-eleven, then “Claude discovered” becomes a rate someone can reproduce. At the moment it is one run that went well, and a report honest enough to say so.

References

  1. Anthropic (2026). Claude discovers a novel enzyme system with CRISPR-like repeats. 23 September 2026. The company’s own announcement.
  2. Yoon, P. H., Athukoralage, J. S., Ameisen, E., Kauderer-Abrams, E., Perry, N. T., & Durrant, M. G. (2026). Autonomous AI agents discover reverse transcriptases with tandem repeat arrays. Anthropic technical report, September 2026. Source of all campaign, rerun, benchmark and expression figures in this essay.
  3. Li, D. B., Brixi, G., Kim, A. S., Fiamenghi, M. B., Driscoll, C. L., Evans, S. A., Gao, A., Ivanova, N. N., Kyrpides, N. C., Deisseroth, K., Fischbach, M. A., & Hie, B. (2026). Coevolutionary mining of prokaryotic non-coding elements with a genome language model. bioRxiv, posted 23 September 2026.
  4. Xin, H., Kitchin, J. R., & Kulik, H. J. (2025). Towards agentic science for advancing scientific discovery. Nature Machine Intelligence 7(9), 1373–1375, 10 September 2025. Background.
  5. Ishino, Y., Shinagawa, H., Makino, K., Amemura, M., & Nakata, A. (1987). Nucleotide sequence of the iap gene, responsible for alkaline phosphatase isozyme conversion in Escherichia coli, and identification of the gene product. Journal of Bacteriology 169(12), 5429–5433, December 1987. Historical background.
  6. Ishino, Y., Krupovic, M., & Forterre, P. (2018). History of CRISPR-Cas from encounter with a mysterious repeated sequence to genome editing technology. Journal of Bacteriology 200(7), April 2018. Historical background.
  7. Jansen, R., van Embden, J. D. A., Gaastra, W., & Schouls, L. M. (2002). Identification of genes that are associated with DNA repeats in prokaryotes. Molecular Microbiology 43(6), 1565–1575, March 2002. Historical background.