Claude AI Discovers a Previously Uncharacterized Phage Reverse-Transcriptase System with CRISPR-Like Repeat Arrays

A general-purpose artificial intelligence system has apparently crossed an important threshold in biological research: not merely summarizing scientific literature or optimizing a human-defined analysis, but noticing an unexpected pattern directly in raw genomic data, deciding that it was scientifically interesting, investigating it further and ultimately directing researchers toward a previously uncharacterized molecular system encoded by bacteriophages. On September 23, Anthropic reported that autonomous Claude agents identified what researchers now call array-associated reverse transcriptases, or ARTs, a family of phage-encoded reverse transcriptase systems associated with unusual arrays of repetitive non-coding DNA. The work has been released as a preprint and has not yet undergone peer review.

Figure 1. Architecture of the ART system and its identification by autonomous genome mining.
(A) Schematic organization of an array-associated reverse transcriptase (ART) locus. ART systems contain an upstream non-coding tandem-repeat array, an ART reverse transcriptase and a neighboring partner protein. Individual repeats are approximately 15–49 nucleotides long and alternate with much longer variable regions of roughly 120–220 nucleotides. (B) Comparison between ART and a conventional CRISPR–Cas system. Although both architectures contain repeated DNA elements separated by variable sequences, ART arrays differ from CRISPR arrays in the length of their variable regions and in the absence of nearby cas genes. ART should therefore not currently be considered a CRISPR–Cas system. (C) Summary of the computational discovery pipeline. Autonomous Claude agents surveyed approximately 1.9 billion protein clusters, identified about 200,000 filtered reverse-transcriptase clusters, examined approximately 11,000 genomic loci and scored 3,564 recurrent neighboring protein families. This analysis ultimately defined 95 ART reverse-transcriptase clusters, 28 of which contained a detectable upstream repeat array. The biological function of ART and its potential programmability remain unknown and require experimental validation.

The study was conducted by Peter H. Yoon, Januka S. Athukoralage, Emmanuel Ameisen, Eric Kauderer-Abrams, Nicholas T. Perry and Matthew G. Durrant at Anthropic. Their technical report, Autonomous AI agents discover reverse transcriptases with tandem repeat arrays, describes both a new biological system and the unusual route by which it was found. The authors deployed multiple instances of Claude Code in an autonomous research harness and asked them to search for previously undescribed systems associated with reverse transcriptases, enzymes that normally copy RNA into DNA.

The scale of the search was substantial. Claude agents surveyed a database containing approximately 1.9 billion protein clusters. They assembled their own hidden Markov model profiles for reverse transcriptases, recovered roughly 200,000 filtered RT clusters, classified them into nine major classes, examined approximately 11,000 genomic loci and scored 3,564 recurrent neighboring protein families as possible functional partners. The complete campaign comprised 119 tasks and 949 agent sessions, consumed 215.6 million tokens and represented 77 agent-hours compressed into 21.5 hours of wall-clock time. According to the authors, this phase proceeded without human intervention.

What makes the discovery especially interesting is that ART was not what the researchers explicitly asked Claude to find. The original mission focused on novel proteins associated with RT genes. During the analysis, one agent encountered a reverse transcriptase associated with jumbo bacteriophages. An initially suspected neighboring protein association was eventually rejected, but the agent decided that the unusual RT deserved further investigation. Because retron reverse transcriptases frequently operate with non-coding RNAs encoded nearby, another agent examined the DNA upstream of related RT genes.

It was while reading this raw sequence directly that Claude noticed an unexpected repeated pattern. The original session transcript records the model recognizing what it described as a tandem repeat array and immediately comparing its architecture with CRISPR arrays, retrons and diversity-generating retroelements. Rather than simply accepting its first interpretation, the agent questioned whether the pattern had already been described, quantitatively characterized the repeats, wrote code to count them and searched the literature for related systems. One representative locus contained 14 copies of a 16-nucleotide repeat separated by unique sequences approximately 100–200 nucleotides long.

That distinction is important because Claude did not discover an entirely unknown reverse transcriptase protein from nothing. Related jumbo-phage RTs had previously appeared in genomic studies. In the MarsHill phage genome, for example, the RT had already been annotated and an upstream non-coding RNA had been proposed. What had apparently escaped recognition was the larger molecular architecture: the RT belonged to a coherent family characterized by a repeated non-coding array, an unusual N-terminal extension and a dedicated neighboring partner protein. Anthropic therefore describes the discovery as recognition of a previously uncharacterized enzyme system rather than the first observation of the underlying RT itself.

Subsequent computational analysis considerably expanded the initial observation. The researchers identified 95 distinct ART RT clusters at 90% sequence identity across cultured jumbo phages and predicted viral sequences. Twenty-eight had a detectable repeat array immediately upstream of the RT. Phylogenetically, the ART reverse transcriptases form a clade close to retrons, but their architecture is different from previously described RT systems. The catalytic YxDD motif characteristic of reverse transcriptases was retained in all 93 complete sequences examined, while ART proteins possessed an unusually long N-terminal region of roughly 180 amino acids before the polymerase domain, compared with around 50 residues or fewer in many related RT families.

The arrays themselves provide one of the strongest reasons for interest. Across the identified ART systems, they range from approximately 0.3 to 4.1 kilobases and contain between three and 21 repeated units. Individual repeats are 15–49 nucleotides long and contain a roughly 15-nucleotide palindromic core. Between them lie much longer, largely unrelated sequences of around 120–220 nucleotides.

At first glance this organization evokes CRISPR, where repeated DNA sequences alternate with variable spacers that ultimately generate guide RNAs. But ART should not be described as a new CRISPR system. Its architecture differs substantially. CRISPR spacers are typically around 30 nucleotides, whereas the variable ART regions can exceed 200 nucleotides. ART units also appear to be conserved between related phages in ways unlike the rapidly changing acquisition pattern of CRISPR spacers, and critically, the researchers found no nearby cas genes. The resemblance is therefore architectural and potentially functional, not evidence that ART is part of the CRISPR-Cas family.

The first experimental evidence that these arrays are biologically active came from an existing RNA-sequencing dataset generated during infection of Staphylococcus lentus by jumbo phage SA1. The researchers found that the ART array is strongly transcribed throughout infection. Fifteen minutes after infection, array-derived RNAs accounted for as much as 8% of all phage RNA, placing them among the most abundant viral transcripts at that time point. Importantly, the signal did not resemble one continuous transcript alone: the RNA resolved into distinct shorter species with reproducible boundaries.

The team then reproduced this behavior experimentally by expressing the SA1 ART locus in Escherichia coli. Small-RNA sequencing again revealed discrete short RNAs derived from the array, both when the native locus was used and when expression was driven from a heterologous promoter. Computational folding suggested that the conserved repeat sequence could form a base-paired stem, consistent with the palindromic character of its core. Together, these observations support the idea that an ART array is processed or expressed as a repertoire of several structurally distinct non-coding RNAs rather than functioning merely as repetitive genomic DNA.

The third component of the system is equally unusual. ART reverse transcriptases are accompanied by dedicated partner proteins, and the authors identified three apparently unrelated partner families. Type I ART systems, representing 59 of the 95 loci, encode a protein of roughly 600 amino acids containing two GCN5-related N-acetyltransferase-like folds. Type II systems, associated with the Staphylococcus phage lineage, encode an approximately 270-residue predominantly helical protein. Type III systems encode a smaller helical protein of roughly 170 residues. The Type II and III partners have no clear homologues in current protein-domain or structural databases.

This three-component organization — non-coding RNA, reverse transcriptase and dedicated partner protein — immediately recalls retrons, bacterial systems in which an RT copies a structured RNA into a characteristic single-stranded DNA molecule and works with an effector that can participate in antiphage defence. ART differs in one potentially important respect: instead of a single RNA element, it appears to encode an entire bank of distinct RNAs. The authors therefore propose a model in which individual ART RNAs might form different complexes with the same RT and partner, potentially allowing one molecular system to respond to multiple molecular signals. For now, this remains a hypothesis.

This is precisely where the distinction between an exciting discovery and a demonstrated biotechnology becomes essential. The researchers have not yet shown that the ART reverse transcriptase actually performs reverse transcription in this system. They have not demonstrated that the array-derived RNAs are substrates of the RT, have not experimentally established direct interaction between the RT and the partner protein, and do not yet know what ART does for the phage. The system could participate in phage competition, host manipulation, defence against other mobile elements or an entirely different process. Its biological function remains open. The authors themselves explicitly state that these questions require experimental characterization.

For the same reason, descriptions of ART as a new programmable gene-editing system would currently be premature. What makes the architecture intriguing is that several other molecular systems containing enzymes paired with libraries of nucleic-acid sequences have eventually proved programmable, CRISPR being the most famous example. Anthropic notes that systems with related combinations of features can perform molecular operations including DNA cutting, copying or insertion. But ART has not yet been shown to perform any of these functions, and architectural resemblance alone cannot establish programmability.

A second major result of the study concerns AI itself. The authors did not merely examine what Claude reported; they investigated why it had noticed the array. When the entire autonomous campaign was repeated ten times, ART-related loci were encountered in most successful searches and two runs investigated the lineage further, but none of the reruns inspected the crucial upstream DNA deeply enough to rediscover the array. The discovery was therefore not deterministic.

The researchers then designed controlled benchmarks in which different Claude models were explicitly given ART sequences. The strongest models were highly effective at recognizing the array when the raw DNA sequence was actually placed in their context. When more elaborate tool-based environments caused models to avoid reading long stretches of DNA directly, performance fell sharply. Across the strongest models, recognition increased substantially as more contiguous DNA was placed into context, reaching as high as 96% for Mythos 5 under one high-context condition.

Anthropic also examined the model's internal activity while it processed the sequence. Two internal Mythos 5 signals were found that responded strongly to the repeated elements. Their activation emerged as the model progressed through successive repeat copies and largely disappeared when the nucleotide sequence of each repeat was shuffled. Immediately after those internal responses, the model produced natural-language reasoning identifying the pattern as a tandem repeat array. The authors interpret this as evidence that the model had developed an internal representation capable of recognizing repeated DNA structures, despite Claude being a general-purpose language model rather than a model trained exclusively on genomes.

That point may ultimately prove almost as consequential as ART itself. Conventional genome-mining pipelines are extremely powerful, but they generally search for patterns specified in advance: a known motif, domain, gene neighborhood or structural relationship. Human scientists remain valuable partly because they can encounter something that does not fit the original search criteria and decide that the anomaly deserves attention. In this experiment, the repeat array lay outside the intended protein-partner search, yet an autonomous agent noticed it, compared it with biological concepts drawn from its broader knowledge, challenged its own interpretation and initiated additional analyses.

Human scientists remained indispensable. They designed the research objective, built the experimental system, reviewed the agent-generated reports and performed all laboratory experiments. Anthropic explicitly states that its wet-lab work is performed by human researchers. The result is therefore better understood as human–AI scientific discovery than as an AI independently operating a laboratory. What is unusual is that the crucial observation and the decision to investigate it were generated inside the autonomous computational campaign rather than being specified beforehand by the researchers.

The finding is particularly relevant to bacteriophage biology because enormous amounts of viral genetic information remain functionally unexplored. Jumbo phages can encode unusually large genomes and molecular machinery that differs substantially from that of their bacterial hosts and smaller phages. Many predicted phage proteins still have no experimentally established function. A system capable of systematically inspecting this sequence space while retaining enough biological context to recognize unexpected architectures could therefore expose molecular mechanisms that conventional annotation pipelines have overlooked.

It could also change the scale of biological discovery. The ART campaign moved from roughly 1.9 billion protein clusters to a small set of candidate systems in less than a day of wall-clock computation. That does not mean that laboratory science has been compressed into 21 hours: determining what ART actually does may require months or years of biochemistry, structural biology, genetics and infection experiments. But it suggests that one of the major bottlenecks in modern biology — deciding which anomalies hidden inside enormous datasets are worth experimental attention — may increasingly be delegated to AI systems.

Feng Zhang, whose work at MIT and the Broad Institute helped establish CRISPR genome editing, reviewed the preprint and described the association between RNA-repeat arrays and reverse transcriptases as “genuinely intriguing,” while emphasizing the need for further investigation. His reaction captures the appropriate interpretation of the result: ART is not yet another CRISPR, and its molecular function remains unknown, but its architecture is unusual enough to justify serious experimental study.

There are therefore two discoveries intertwined in this report. The first is biological: a previously uncharacterized family of jumbo-phage reverse-transcriptase systems combining a repetitive RNA-producing array, an unusual RT and a dedicated partner protein. The second is methodological: a general-purpose AI model encountered an anomaly in primary genomic data that was not part of its original search target and pursued it far enough to produce a testable biological discovery.

Whether ART eventually becomes a programmable biotechnology, reveals a new mechanism of phage biology or proves to perform a much narrower natural function cannot yet be known. But even the current evidence is significant. The study shows a plausible route by which autonomous AI agents can move beyond literature synthesis and predefined data analysis into the earliest stage of scientific discovery itself: seeing something unexpected and recognizing that it may matter.





Sources :

Peter H. Yoon, Januka S. Athukoralage, Emmanuel Ameisen, Eric Kauderer-Abrams, Nicholas T. Perry & Matthew G. Durrant. Autonomous AI agents discover reverse transcriptases with tandem repeat arrays. Anthropic preprint, 2026 : https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf

Anthropic — Claude discovers a novel enzyme system with CRISPR-like repeats, September 23, 2026 : https://www.anthropic.com/news/claude-discovers-novel-enzyme-system

Comments

Most Consulted Articles

Bacteriophages Disarm Inflammation-Associated E. coli Without Erasing the Gut Microbiome in IBD

History Part 12 : Post-War Stagnation and Phage Therapy’s Marginalization in the West (1945–1980s)

🌐 PhageAtLabs® — a global interactive mapping of academic laboratories in phage research.

The Phage Therapy in the spotlight !

Groundbreaking achievement : Phagos raises €25m to end bacterial disease