Sixteen bacteriophages designed by a generative model infected and killed E. coli. The story has been told as if the danger were the machine that writes genomes. It isn’t. The danger is that the lock protecting us was on a different door, and it has spent years recognising faces.
Sixteen clear spots on a Petri dish
On 6 August 2026, Science published “Generative design of bacteriophages with genome language models”, by Samuel King, Brian Hie and their team at the Arc Institute and Stanford. The headline that travelled was this: AI has created living viruses for the first time.
That’s true. And like almost everything true, it’s been badly told.
Underneath the headline there’s a funnel. The team used Evo 1 and Evo 2, two genome language models — the same idea as an LLM, but trained on DNA instead of text — and asked them to write complete genomes in one pass. Not a gene. Not a protein. The whole genome, start to finish.
Thousands of designs came out. (A hygiene note here: much of the coverage speaks of some 700,000 generated genomes; the paper and the preprint say “thousands”. I’ll go with the primary source.) A computational filtering framework cut that to 302 candidates. Chemical synthesis worked for 285. And when they were dropped onto E. coli cultures, 16 produced that unspectacular, utterly conclusive thing: a clear spot on the plate. Dead bacteria. A 5.6% hit rate on what was synthesised. And three of them — Evo-Φ69, Evo-Φ100 and Evo-Φ111 — finished ahead of the original phage in head-to-head competition experiments.
The model wasn’t starting from absolute zero. It was seeded with conserved fragments of ΦX174, a tiny bacteriophage of about 5,400 nucleotides of single-stranded DNA that knows how to do exactly one thing: get inside E. coli and burst it from within. The models were then fine-tuned on some 15,000 Microviridae genomes to improve the aim.
Choosing ΦX174 wasn’t an accident or marketing. It’s the most handled virus in the history of biology. It was the first DNA genome ever fully sequenced, by Frederick Sanger in 1977. And in 2003 Craig Venter and Hamilton Smith’s team rebuilt it in full from synthetic oligonucleotides in fourteen days. When you want to show a new technique works, you test it on the best-mapped terrain you have. That’s good science.
The uncomfortable 97%
This is where the volume should come down, because there’s a sceptical reading that holds up well.
Oliver Crook, a computational biochemist at Oxford, ran his own analysis and measured how close those 16 phages were to their template. The answer: 97% average identity with ΦX174. His line is a good one: “what we saw, at a very plain view, were brothers and sisters of the original virus. They’re not fundamentally behaving in a new way.” They fell within the spread of diversity that already exists among natural phages.
In the same vein, Jordi García Ojalvo (Pompeu Fabra) points to the low efficiency of the process: thousands of genomes generated, 16 viable phages out. Heather Hendrickson (University of Canterbury) adds the more uncomfortable part: “an AI can make a genome that makes a passable version of a phage, but that machine does a poor job of telling us what it ‘learned’”. And Simon Jackson (Waikato) sums it up without ornament: around 5% of the designs worked; this is an early proof of concept. One detail Jackson flags that I love: half the functional phages had picked up mutations, meaning natural evolution lent a hand polishing the AI’s designs.
The Stanford team has a reply, and it also holds up. King argues that sequence identity doesn’t tell the whole story: several designs differed in three-dimensional protein structure, in growth kinetics and in infection dynamics. And there’s one detail I find the most interesting in the entire paper: one phage, Evo-Φ36, swapped gene J — the one that packages the genome and holds the capsid together — for the homologue from the evolutionarily distant phage G4. That precise swap had been tried before and was known not to be viable. Here it was. The model saw a context-dependent compatibility we hadn’t.
My honest read: life did not appear out of nothing. Something more modest and more important did. A system that can navigate the space of viable genomes with its own judgement, and gets it right often enough for the lab to confirm. As Víctor de Lorenzo (CSIC) puts it, natural evolution has explored only a tiny fraction of the possible functional space. Now there’s a machine sampling that space without waiting for a generation to pass.
Why this isn’t about viruses, it’s about bottlenecks
And here’s where I want to take the article.
In ecology, a bottleneck is the moment a population gets squeezed: an event cuts the number of individuals and, with it, all the variability that was there. What survives isn’t the best, it’s whatever happened to be on the right side of the funnel. The shape of the species for the next thousand years is decided by that squeeze, not by anyone’s talent.
Biosecurity works the same way. Secrecy has never been what protected us. What protected us is a chain of chokepoints, and only one of them has to hold:
- Knowing what to write. The sequence that does harm.
- Getting someone to manufacture it. Synthetic DNA doesn’t come out of a drawer.
- Assembling and booting it. Getting that molecule into a cell and making the thing wake up.
- Tacit skill. Hands, cultures, contamination, the 90% that isn’t written in any protocol.
For decades, the real chokepoint was the first one. Writing a genome that works was ruinously expensive in expert knowledge. And that’s the August news: that link just loosened. It hasn’t vanished — a 5.6% success rate shows there’s still friction — but it loosened, and the Evo 2 weights are published and downloadable by anyone.
When a bottleneck loosens, the right question isn’t “how frightening, the machine writes.” It’s: which link is the shortest one now? Because a barrel holds what its lowest stave allows, always.
The answer is link two. DNA synthesis screening. It’s the one point in the whole chain where somebody, somewhere, looks at what you ordered before making it for you.
And that lock has a beautiful design problem.
An immune system that only recognises familiar faces
Here’s how screening works today: when you order a sequence from a commercial provider, software compares it against databases of known pathogens and toxins. If it looks close enough to something dangerous, it flags. This is homology-based screening: recognition by resemblance.
If that sounds familiar, it’s because it is precisely our innate immunity. Our first line of defence doesn’t understand intentions; it recognises molecular patterns burned in at the factory. It works beautifully against what it has been seeing for a hundred million years, and it is blind to anything that doesn’t resemble the catalogue.
On 2 October 2025, a consortium led out of Microsoft Research — first author Bruce Wittmann, Eric Horvitz as senior author, with Twist Bioscience, IBBIS, Battelle and the University of Birmingham on board — published exactly that blindness in Science. They used generative protein design tools to redesign known toxins: same toxic function, different sequence. Then they ran them through the screening systems. Many got through. They called it, quite deliberately, a biological “zero-day”: an unknown vulnerability in the security infrastructure, the same kind that shows up in software.
They behaved like good security researchers. They warned the US government and the vendors before publishing, patches went out, they didn’t release the full code or name the toxins. But the write-up admits the essential part: even after patching, some variants still escape detection.
There it is, in one sentence. The key has changed shape and the lock is still looking at the face.
And there are two more holes in the same wall. First: screening has been, for most of this time, voluntary. The OSTP framework for nucleic acid synthesis screening (April 2024) effectively binds only those who depend on US federal funding; the rest rests on the good practice of the International Gene Synthesis Consortium, which is a club, not a law. And a May 2025 executive order left that framework under review, pending replacement. The second hole, newer: benchtop synthesisers. Machines that make DNA in your own lab, with no order, no provider and nobody looking. Regulation controls them at the point of sale and loses sight of them afterwards.
The law is coming, and it’s looking at the wrong place
To be fair, governance has moved. Two serious things happened in 2026.
One: on 4 June, Sam Altman (OpenAI), Dario Amodei (Anthropic), Demis Hassabis (Google DeepMind) and Mustafa Suleyman (Microsoft AI) — people who don’t usually sign anything together — published an open letter asking the US Congress to make synthetic DNA screening mandatory, along with recordkeeping on orders. When four competitors ask to have their own substrate regulated, it’s worth listening.
Two: there’s a bill, the Biosecurity Modernization and Innovation Act of 2026 (S.3741), introduced on 29 January 2026 by Tom Cotton with Amy Klobuchar. It does sensible things: makes screening mandatory for covered providers and supplants the voluntary guidance, brings benchtop synthesiser vendors into the definition, contemplates a mechanism for detecting orders split across several providers, and sets civil penalties up to $500,000 for individuals and $750,000 for organisations.
And still, the critical analysis of the text puts a finger on the point that matters: the law mandates screening, but doesn’t change how screening works. It’s still homology. Still recognising faces. The vulnerability Microsoft demonstrated in 2025 remains intact, now with an official seal and a fine for not applying it. On top of that, split-order detection is drafted as “may” rather than “shall”, which in a legal text is the difference between a mechanism and an intention.
That’s what I wanted to get at with the bottleneck. We’re fitting a new lock, more expensive and better documented, to the door we already knew how to open.
The bright side, which is enormous and is not a consolation
None of the above invalidates why this work was done.
Antibiotic resistance kills. The GRAM project’s global analysis published in The Lancet in September 2024 estimates more than 39 million deaths directly attributable to resistant infections between 2025 and 2050, and counts over a million deaths every year between 1990 and 2021. That isn’t a scenario. It’s the baseline.
Phage therapy has been waiting a century for its moment. Its big practical problem was always the same as with antibiotics: the bacterium learns. And here’s the result from the Stanford paper that actually matters, beyond the headline: a cocktail of the generated phages overcame resistance in three E. coli strains within one to five passages, while ΦX174 on its own failed completely. Hie says it plainly: “if the bacteria gain resistance to a single phage, it’s game over for the medication”. The point isn’t designing a virus. It’s being able to design a hundred and rotate them faster than the bacterium adapts.
And the team did do its safety homework, which deserves saying: they excluded viruses that infect humans and animals from training, worked with a phage that only attacks E. coli, and followed a strict protocol that Juli Peretó (Universitat de València) explicitly credits. Hie further argues that an AI design is less risky than a natural pathogen, because you can build safety checks into the design process, and you can’t ask evolution for that.
All of that is true. And still, Simon Clarke (Reading) presses where it hurts: there is no guarantee that everyone attempting something similar will be equally careful.
The backdrop I can’t shake off
I’ll leave you with the asymmetry, which is the dystopian part, and I don’t feel like dressing it up.
Defence has to work every time. Attack, once. In computer security we accepted that decades ago, which is why we built defence in depth, coordinated disclosure and patches. In biosecurity we have one layer — homology screening — it’s voluntary across half the world, it was shown to be evadable less than a year ago, and the law coming to fix it doesn’t touch the evasion.
I’d put it like this: the ability to compose viral genomes with generative AI already exists; the governance to steer it safely does not.
And there’s no way back through closing off knowledge. Evo 2 is downloaded onto thousands of hard drives. That debate is over, whatever anyone’s view on open weights. What is still open is the other part, which also happens to be the cheap part:
- Screen by function, not by resemblance. Predict what a sequence does, not what it looks like. That’s an AI problem, and it so happens we have AI.
- Mandatory and global. A lock fitted in one country only is a list of addresses where there is no lock.
- Put screening inside benchtop synthesisers. If the device prints DNA on your desk, the check has to live in the firmware, not in the order form.
- Change “may” to “shall” in split-order detection. It costs one word.
Four things. None of them needs a better model. All of them need someone to decide who answers for this.
And that, almost always, is the part that doesn’t get automated.
Ideas over code; evidence over smoke.

Leave a Reply