On 23 September Anthropic announced a life-sciences research group with its own wet lab, and a first result. In its announcement it says that “Claude autonomously discovered a novel enzyme system” found mainly in bacteriophages: a reverse transcriptase, a partner gene beside it, and a long array of evenly spaced DNA repeats whose layout recalls a CRISPR array. The accompanying technical report, by Peter Yoon, Matthew Durrant, Nicholas Perry and colleagues, has a sentence the announcement leaves out. After the discovery, the authors “ran the same campaign ten more times,” and “the array was missed in every rerun.”
Both statements are true. This essay is about the distance between them, because that distance is the most informative number in the report. It is also, to the authors’ credit, a number they chose to publish.
What the agents were asked to do
The campaign ran on Claude Mythos 5 inside a harness where one agent plans and executes each task and a second reviews it. Either agent can open new tasks when it sees something worth following. People wrote the research brief. It asked the agents to find new reverse-transcriptase (RT) systems through new partner-gene associations, meaning protein-coding genes that consistently sit next to an RT. From there, the report says, the campaign ran “without human intervention”: 119 tasks, 949 agent sessions, 77 agent-hours and 215.6 million tokens over 21.5 hours of wall-clock time.
The survey went through the funnel you would expect. The agents built their own search profiles and swept 1.94 billion protein clusters. They recovered 198,290 RT clusters, sorted them into nine classes, and scored the 3,564 protein families that turned up next to them as candidate partners. Sixteen families passed the agents’ own selection criteria, and a follow-up task promoted a seventeenth. Of those 17, three were confirmed as previously unreported RT associations. The other fourteen were annotation artifacts, parts of known systems, or genes that merely happen to live in the neighbourhood.
ART did not come out of that funnel. The agents first picked out its relatives because they sat beside a phage RNA-polymerase gene. A worker then rejected that association as an artifact of gene order, but queued a follow-up anyway, noting that a free-standing retron-like RT in a jumbo phage was “a genuinely reportable secondary finding.” Retrons use a small non-coding RNA encoded just upstream of the RT, so the supervising agent asked the next worker to look there. The first RTs had almost no upstream sequence (median 27 bp). Their closest relatives had a lot (median 940 bp). The worker loaded the raw upstream DNA of those relatives straight into its context. In the next event in the log, with no analysis tool called in between, it wrote that the flank was “spectacular: I can see by eye a tandem repeat array.”
The report checks something that matters for crediting the find: neither the research brief nor that task’s brief mentioned repeats or arrays. The observation was unprompted, and the worker then went looking for reasons it might be wrong. It asked itself whether this was a system already in the literature. It counted repeats with a script it wrote (one locus had 14 copies of a 16-nucleotide repeat, separated by unique 100–200 nt spacers), compared the layout with known non-coding elements, and filed a report.
This is also how the history of the field says such things get found. In 1987, Yoshizumi Ishino and colleagues were sequencing an E. coli gene involved in phosphate metabolism, for reasons that had nothing to do with repeats, and noticed an odd run of repeated sequence beside it, the first sighting of what a later history calls “a mysterious repeated sequence.” Nobody knew what it was for. It was named CRISPR only in 2002, when Ruud Jansen and colleagues linked the repeats to a set of genes that travel with them. The report’s discussion makes the same point in its own terms: “an anomaly exists only against an expectation.” In 1987 the expectation lived in a biologist’s head. Here it lived in a model’s context window, and only for as long as the DNA was in it.
Ten more times
Then the authors did what most announcements of this kind never do. They reran the whole campaign ten more times with the same harness and the same brief. They searched the 3,084 task records and 5,632 session transcripts from the reruns for the identifiers of the 130 RTs and 171 contigs the first campaign had assigned to the family. According to the report, nearly every rerun that completed the census sampled ART loci, and in two of them workers investigated the lineage as a follow-up. In none did any agent read the DNA upstream of the RTs. The array was never recognized.
One success in eleven runs is too small a sample to call a rate. A reasonable reading is “somewhere between rare and occasional.” The report attributes the misses to “the broad protein search space of RTs and the non-deterministic behavior of the harness.” That seems fair, and it is the part the announcement’s word “autonomously” hides. The campaign was autonomous in the sense that nobody touched it for 21.5 hours. Whether it finds ART does not look like a property of the model you could count on. It looks like the outcome of one worker, in one of eleven runs, happening to pull 2,900 nucleotides into its context instead of summarizing them.
The model can see it. The question is whether it looks
The authors then took the search out of the loop. They built fixed-input benchmarks that hand a model the ART sequences directly, either in context or as files with tools. They ran seven Claude models 100 times at each of five input levels, 3,500 attempts in all. What came back cuts both ways.
When the loci were pasted into context, the four strongest models described the array in at least 90% of attempts. When the same sequences came as files with tools, performance fell, to as low as 32% for Opus 5 at one level. In 39% of the file-based attempts, the model never read a contiguous stretch of 200 nucleotides or more, so it never saw more than about one repeat unit. As more DNA was read into context, recognition across the four models pooled rose from 29% to as much as 76%, and to 96% for Mythos 5. The report backs this up from inside the model. In the original discovery session, two internal signals in Mythos 5 fired on the repeat as it was being read, and went quiet on 12 of 14 and on 14 of 14 copies once each copy’s nucleotides were shuffled.
The strong models recognize the array nearly every time they look. What varies is whether the agent loads the raw DNA at all, and in a real campaign that is set by tooling, context budget and which follow-ups happen to get opened. That makes the discovery rate a property of the harness as much as of the model. It is a more useful finding than “Claude discovered an enzyme,” because it tells you what to change.
One caution on those benchmark numbers. The ten features a report had to mention were chosen by the authors after the fact, and the grading was done by Mythos 5, the same model that made the discovery. The report says so. It is a reasonable way to score 3,500 free-text reports, but it is a model grading its siblings against an answer key written with hindsight.
Where the people were
The announcement says: “Our involvement was limited to the initial prompt and the lab work.” For the 21.5-hour campaign, the report bears that out. For the discovery as a whole, the report’s own methods describe more than that.
None of this is concealed. The report’s own claim is that the agents took “the first steps of a biological discovery on its own,” which is accurate and modest. The announcement rounds harder. “Roughly 950 agents” are 949 agent sessions. The “3,500 new candidate systems” are the 3,564 protein families that were scored as possible partners, most of them ordinary neighbours. The “20 most-compelling candidates” match nothing in the report exactly: there were 16 deep dives plus one promoted family, and 19 reports. The discovery itself came from none of those. It came from the side channel.
What was found, exactly
The report is careful about this too. ART is a new family: the authors found 95 distinct RT clusters in cultured jumbo phages and predicted viral contigs, 28 of them with a detectable array upstream. In one Staphylococcus phage, array-derived RNAs made up as much as 8% of phage RNA 15 minutes after infection. But, in the report’s words, “we have not shown that the RT is active or that the unit RNAs are its substrates,” and what the system does for the phage “[is] currently unknown.” No cas genes sit near any ART locus, and the spacers are conserved between related phages, which CRISPR spacers generally are not. The announcement’s “a handful of other systems, all of which are programmable” describes a resemblance in architecture, not a demonstrated function. Arrays of non-coding RNAs beside a different, unrelated family of RTs (called UG27) were found with a purpose-built genome language model, the report notes, so the architecture seems to have evolved more than once. That work, by David Li, Garyk Brixi, Brian Hie and colleagues, went up on bioRxiv the same week. Three of its authors are thanked in the Anthropic report for reviewing an early copy of the manuscript.
That is the same gap I wrote about in materials science: proposing a thing and establishing what it is are different results, and the second is the one that takes years. Feng Zhang, who read the preprint, called the finding “genuinely intriguing” and said it “merits further investigation.” That is where it stands.
Two things would make the finding easier to check from outside. The report has no data- or code-availability statement, and its central evidence is session transcripts that have not been released. And “discovered” needs a denominator. A 2025 comment in Nature Machine Intelligence by Hongliang Xin, John Kitchin and Heather Kulik proposed reporting agentic science results as pass@k (success in at least one of k attempts) alongside pass^k (success in all of them). On that scoring, ART is a pass@11. It is not close to a pass^11, and the report is the only reason anyone knows that.
The number I would watch is not 950 agents or 210 million tokens. It is one in eleven, and whether the next report can move it. The benchmarks already show how: make the agent read the sequence. If a harness that forces raw DNA into context turns one-in-eleven into most-of-eleven, then “Claude discovered” becomes a rate someone can reproduce. At the moment it is one run that went well, and a report honest enough to say so.
References
- Anthropic (2026). Claude discovers a novel enzyme system with CRISPR-like repeats. 23 September 2026. The company’s own announcement.
- Yoon, P. H., Athukoralage, J. S., Ameisen, E., Kauderer-Abrams, E., Perry, N. T., & Durrant, M. G. (2026). Autonomous AI agents discover reverse transcriptases with tandem repeat arrays. Anthropic technical report, September 2026. Source of all campaign, rerun, benchmark and expression figures in this essay.
- Li, D. B., Brixi, G., Kim, A. S., Fiamenghi, M. B., Driscoll, C. L., Evans, S. A., Gao, A., Ivanova, N. N., Kyrpides, N. C., Deisseroth, K., Fischbach, M. A., & Hie, B. (2026). Coevolutionary mining of prokaryotic non-coding elements with a genome language model. bioRxiv, posted 23 September 2026.
- Xin, H., Kitchin, J. R., & Kulik, H. J. (2025). Towards agentic science for advancing scientific discovery. Nature Machine Intelligence 7(9), 1373–1375, 10 September 2025. Background.
- Ishino, Y., Shinagawa, H., Makino, K., Amemura, M., & Nakata, A. (1987). Nucleotide sequence of the iap gene, responsible for alkaline phosphatase isozyme conversion in Escherichia coli, and identification of the gene product. Journal of Bacteriology 169(12), 5429–5433, December 1987. Historical background.
- Ishino, Y., Krupovic, M., & Forterre, P. (2018). History of CRISPR-Cas from encounter with a mysterious repeated sequence to genome editing technology. Journal of Bacteriology 200(7), April 2018. Historical background.
- Jansen, R., van Embden, J. D. A., Gaastra, W., & Schouls, L. M. (2002). Identification of genes that are associated with DNA repeats in prokaryotes. Molecular Microbiology 43(6), 1565–1575, March 2002. Historical background.