Nature Microbiology Study Uses Machine Learning to Predict Which Phages Will Infect Which Bacterial Strains
One of the fundamental problems in phage therapy is deceptively simple: given a patient's bacterial isolate and a collection containing hundreds or thousands of bacteriophages, which phage should be tested first?
A study published in Nature Microbiology presents a machine-learning framework designed to answer that question directly from bacterial and phage genome sequences. Rather than relying on a predefined receptor, a particular bacterial lineage or previous knowledge of the molecular mechanisms controlling infection, the system attempts to predict whether an individual phage will infect an individual bacterial strain.
The work was led by Avery J. C. Noonan and colleagues from Lawrence Berkeley National Laboratory, the University of California Berkeley and Pennsylvania State University, with Vivek K. Mutalik and Adam P. Arkin among the senior authors. The study brings together computational modelling, large phage–host interaction datasets and experimental genetic validation to address a problem that becomes increasingly important as phage banks grow.
The specificity of bacteriophages is one of their most attractive properties therapeutically. Unlike many broad-spectrum antibiotics, a phage can potentially target a particular bacterial population while leaving much of the surrounding microbiota unaffected. But this precision creates a logistical problem: two strains belonging to the same bacterial species can respond very differently to the same phage.
Phage susceptibility depends on far more than taxonomy. Bacterial receptors, surface structures, capsules, lipopolysaccharides, restriction systems, defence mechanisms and mobile genetic elements can all influence whether infection succeeds. A phage that kills one Escherichia coli strain may therefore fail completely against another.
This becomes increasingly difficult to manage experimentally when phage collections contain hundreds or thousands of candidates. Testing every possible phage against every newly isolated bacterial strain is possible in principle, but it becomes expensive, slow and difficult to scale.
The researchers therefore developed what they describe as a phylogeny-agnostic machine-learning framework. The objective was not simply to predict the bacterial species associated with a phage, as many host-classification tools do, but to predict infection at the level of individual phage–strain pairs.
The scale of the computational optimization was substantial.
The team assembled six previously published datasets containing a combined 115,037 experimentally measured phage–host interactions involving 949 bacterial strains and 518 phages. They then evaluated more than 200 modelling configurations and trained more than 13.2 million individual predictive models before selecting the final workflow.
Instead of requiring predefined infection genes, the model represents bacterial and phage genomes according to the presence or absence of protein families. Machine-learning algorithms then identify genomic features associated with successful or unsuccessful infection.
The final system uses recursive feature elimination to identify informative genomic variables and CatBoost gradient-boosted decision trees to generate predictions. Importantly, the model does not begin with assumptions such as “this tail fibre recognizes this receptor” or “this bacterial capsule determines susceptibility.”
It instead attempts to learn these relationships from large collections of genomes paired with experimental infection outcomes.
Across the different datasets, performance ranged from an AUROC of 0.67 to 0.94 when predicting whether known phages could infect previously unseen bacterial strains. For E. coli, the model reached an AUROC of approximately 0.87, comparable to a previously reported method that relied on known E. coli-specific determinants of phage infection.
That distinction is important. Species-specific models can perform very well when the relevant biology is already understood. But they become less transferable when researchers move toward less characterized bacteria or phages.
The new framework was designed to reduce that dependence on prior biological knowledge.
The researchers tested several situations that resemble practical phage selection. One involved predicting whether phages already present in a collection could infect a previously unseen bacterial strain. Another examined a known bacterial strain confronted with a previously unseen phage. The most difficult scenario involved both a new bacterial strain and a new phage.
Performance was strongly dependent on the quantity and quality of available interaction data. Larger datasets consistently produced more reliable predictions. One of the weakest-performing datasets contained only 3,658 phage–host interactions and fewer than 5% positive interactions, highlighting a major difficulty for machine learning in this field: most phage–bacterium combinations do not result in productive infection.
This makes phage datasets exceptionally imbalanced.
In some published collections used in the study, only a few percent of tested interactions were positive. A model therefore has to learn meaningful biological signals from a very large background of failed infections without simply learning to predict “no infection” for almost every pair.
The researchers attempted to address this through class weighting, feature selection and extensive cross-validation.
But computational performance alone was not enough. A model could theoretically achieve convincing results within previously collected datasets while failing completely when confronted with phages generated elsewhere.
The team therefore performed an independent experimental validation.
They tested 52 previously unseen phages from the BASEL collection against 25 E. coli strains. This generated a matrix of 1,300 experimentally tested phage–host combinations. Sixty produced unclear phenotypes, leaving 1,240 interactions that could be directly compared with the computational predictions.
The model achieved an AUROC of 0.84.
This was slightly below its internal cross-validation performance of approximately 0.87 for E. coli, but the result is important because the tested phages had not been part of the original training set and the experimental data were generated independently.
The study therefore provides evidence that the model was not simply memorizing the phages on which it had been trained.
The researchers then asked another important question: was the model learning biologically meaningful infection mechanisms, or merely finding statistical correlations?
To investigate this, they used random barcode transposon-site sequencing, or RB-TnSeq, in E. coli ECOR27. This approach involves creating large bacterial mutant libraries in which individual genes are disrupted and then challenging those mutants with different phages.
If deleting a particular bacterial gene dramatically changes survival during phage exposure, that gene is likely involved in infection or resistance.
The researchers screened 3,804 E. coli gene knockouts against 19 bacteriophages and identified 51 high-scoring genes associated with phage infection.
They then compared those experimentally identified determinants with the genomic features independently selected by the machine-learning model.
Thirty-five of the 51 genes, or 68.6%, could be linked to predictive features identified computationally, either directly, through nearby genes or through known functional relationships.
This provides an important degree of biological validation.
The model recovered signals associated with several well-established mechanisms of phage infection, including outer-membrane proteins, lipopolysaccharide and O-antigen biosynthesis, capsule-associated pathways and bacterial defence systems.
Among the predictive features were genes connected to porins such as Tsx, FadL and NfrA, as well as components involved in bacterial surface architecture. Defence systems generally decreased predicted infection probability, while anti-defence factors were associated with increased probability.
The model also identified features whose connection to phage biology is not yet clear.
Among the 25 strongest predictive signals in the E. coli dataset, eight lacked an obvious known mechanism linking them to phage infection. The researchers highlight potential associations involving putrescine metabolism and regulation surrounding the known phage receptor FepA as examples requiring additional experimental investigation.
This is potentially one of the most interesting aspects of the approach.
A predictive model that only rediscovers known receptors would essentially automate existing knowledge. A model capable of repeatedly identifying unexplored genetic determinants could instead become a tool for discovering new phage–host biology.
The researchers then moved from prediction to a more directly therapeutic problem: designing phage cocktails.
A common strategy in phage therapy is to combine multiple phages, increasing the probability that at least one component will infect the target bacterium and potentially making it more difficult for bacteria to evolve resistance to the entire preparation.
But selecting cocktail components can itself be empirical.
The team developed a workflow that grouped phages according to their predicted functional profiles and then selected high-probability candidates from different groups. This was intended to combine predicted activity with mechanistic diversity rather than simply choosing the phages that infected the greatest number of strains in previous experiments.
The model-guided approach was compared with a simpler strategy based on selecting broadly infectious, or “promiscuous,” phages.
For single-phage selection, the computational approach identified active phages for approximately 66.9% of unseen E. coli strains, 66.7% of Klebsiella strains and 39.4% of Vibrionaceae strains.
The improvement was particularly marked for Klebsiella, where single-phage selection performed up to 3.1 times better than the promiscuity-based approach.
Increasing the cocktail to five phages substantially expanded predicted coverage.
Five-phage cocktails selected by the model contained at least one active phage for 97.5% of E. coli strains, 87.8% of Klebsiella strains and 57.5% of Vibrionaceae strains in the corresponding analyses.
These numbers should not be interpreted as demonstrating a 97.5% clinical success rate.
The experiments measure whether a selected cocktail includes a phage predicted or experimentally shown to interact with bacterial strains in the datasets. Successful treatment in a patient involves many additional factors, including pharmacokinetics, immune clearance, bacterial location, biofilms, phage resistance, formulation, dose and interactions with antibiotics.
Nevertheless, the result demonstrates how computational screening could change the order in which phages are tested.
Instead of screening an entire bank experimentally, clinicians or researchers could theoretically use genomic information to generate a ranked list of the phages most likely to infect a patient's bacterial isolate and then concentrate laboratory susceptibility testing on those candidates.
That would not eliminate the need for a phagogram.
It could make the phagogram more efficient.
The distinction is important because the model is not sufficiently accurate to justify selecting a therapeutic phage without experimental confirmation. An AUROC of 0.84 represents useful predictive enrichment, not certainty at the level of an individual infection.
The study also identifies clear limitations.
Despite being described as phylogeny-agnostic, the framework does not mean that a model trained on one bacterial genus can currently predict host range equally well in completely unrelated genera.
In fact, cross-genus transfer remained difficult.
When models were trained while withholding an entire bacterial genus, performance fell close to random levels in several tests, with AUROC values around 0.55–0.60. Combining unrelated datasets also failed to provide the improvements observed when related Klebsiella datasets were combined.
This suggests that many determinants of phage susceptibility remain genus-specific or are represented differently enough that the current protein-family approach cannot capture universal rules of infection.
The phrase phylogeny-agnostic therefore refers primarily to the architecture of the method: the framework does not require predetermined phylogenetic rules or known receptor mechanisms. It does not imply that phage–host biology itself has become independent of evolutionary relationships.
The study also highlights another limitation that affects the entire field: phage susceptibility datasets are not standardized.
Some of the datasets analysed were generated using solid agar assays, whereas others relied on liquid-culture measurements. Different research groups also use different thresholds to classify an interaction as positive or negative.
Those methodological differences can introduce noise when multiple datasets are combined.
The authors therefore argue that larger, standardized and community-generated interaction datasets could substantially improve future models.
This could create an important feedback loop.
Experiments would generate larger phage–host matrices. Machine learning would identify the most informative untested combinations. Those experiments would then be performed and incorporated back into the model, progressively improving prediction while reducing unnecessary screening.
The researchers specifically propose active-learning approaches in which the model's uncertainty is used to decide which experiments should be performed next.
Future systems could also incorporate information beyond simple protein-family presence or absence. The authors point toward representations capable of capturing structural information, including protein language-model embeddings, as one potential route toward better transfer between different bacterial groups.
For phage therapy, the broader importance of the study lies in scalability.
The growth of phage banks creates a paradox. Larger collections increase the probability that a matching phage exists, but simultaneously make exhaustive susceptibility testing increasingly difficult.
Genome-based prediction could provide an intermediate layer between sequencing and wet-lab testing.
A patient's bacterial isolate could be sequenced, its genome compared computationally with a characterized phage collection, and the highest-probability candidates prioritized for experimental confirmation.
This would preserve the biological specificity that makes phages attractive while potentially reducing the time required to identify a suitable treatment.
The same technology could extend beyond medicine.
Because the framework predicts strain-level interactions rather than merely assigning a broad taxonomic host, the authors envision applications in precision microbiome engineering, agriculture and industrial microbiology, where researchers may want to selectively remove particular bacterial populations without disrupting entire microbial communities.
The study therefore does not replace experimental phage susceptibility testing, nor does it solve the host-range problem across all bacterial species.
What it demonstrates is that enough information is contained within bacterial and phage genomes to make strain-level infection meaningfully predictable.
That represents a notable shift.
Phage matching has traditionally been treated primarily as an experimental problem: expose bacteria to phages and observe which ones work. The Nature Microbiology study suggests that it can increasingly become a combined computational and experimental process, where sequencing narrows the search before the first plaque assay is performed.
As phage collections continue to expand and personalized therapy requires faster matching between bacterial isolates and therapeutic viruses, that ability to move from thousands of possible phages to a manageable shortlist could become one of the most practical applications of machine learning in phage therapy.
Sources
Avery J. C. Noonan, Lucas Moriniere, Edwin O. Rivera-López, Krish Patel, Melina Pena, Madeline Svab, Alexey Kazakov, Adam Deutschbauer, Edward G. Dudley, Vivek K. Mutalik & Adam P. Arkin. Phylogeny-agnostic strain-level prediction of phage–host interactions from genomes using machine learning. Nature Microbiology (2026). s41564-026-02482-5 (1)

Comments
Post a Comment