Over the past few days, I have watched several scientists and biotech experts on X react with real frustration to claims about the bio-risks of advanced AI, largely spurred by the threat report released by Anthropic. And although I disagree with some of their conclusions, I think people in EA -- working on AI safety and biosecurity -- should pay closer attention to what is producing that reaction.
The discussion began, at least for me, with a thread by by a scientist who said he had experience training a LLM and synthesizing viruses in a lab. His argument, stated bluntly, was that many of the claims about how AI can be used to create dangerous viruses are detached from the reality of biology.
Soon afterward, another scientist who has also worked on frontier AI and physically made viruses in a lab, made a similar argument, although he focused more on what happens after a virus has been engineered and released. I think some of his conclusions go further than the evidence supports, but the fact that scientists with relevant experience see the current discussion as badly miscalibrated makes their underlying concern harder to wave.
I do not share this conclusion that AI has little effect on bio-risk. I do think parts of the AI-bio risk community have a communication problem: we sometimes move too quickly from what a model can produce on a screen to what an actor can accomplish in the physical world, while some critics make the opposite mistake, treating the continued existence of physical bottlenecks as evidence that AI cannot significantly change the overall pathway.
Some of the disagreement is substantive, and some of it comes from people talking about very different scenarios as though they were debating the same claim.
A quick caveat: I am not trying to estimate the probability of an AI-enabled bio catastrophe here, and I am especially not arguing that current evidence rules one either ir or out. My narrower claim is that we often communicate the risk at a level of abstraction that conceals many of the assumptions on which our conclusions depend.
Depending on who is speaking, this might mean that a model can generate a viral sequence, propose modifications to an existing pathogen, predict some of their effects, help a scientist select candidates, or participate in a longer process that eventually produces a pathogen with particular properties. These capabilities are related, but not equivalent.
Between a model producing an answer and an actor causing biological harm, someone generally needs to evaluate the output, obtain the relevant materials, assemble them correctly, recover a viable biological system, test it, interpret the results, respond to failures and repeat parts of the process. Even after someone successfully produces a viable virus, the story is not finished, because the relevant traits must survive replication, behave as expected in human hosts, and persist during the transmission long enough to produce harm at the scale being discussed.
When these steps are compressed into the claims that "AI can design a dangerous virus", a biologist may reasonably hear that computational design and realized biological capability are being treated as roughly the same thing. The person making the claim may have a much richer threat model in mind, but if the intermediate assumptions remain unstated, the disagreement is hardly surprising.
In fact, this problem was pretty evident in the responses to the thread on X. Some people were discussing a future AI system that could control sophisticated cloud labs, robotics and other parts of the economy; others were describing a nearer-term system that might persuade or pay humans to perform the physical work; the OP was largely responding to claims about what present or near-term systems could accomplish.
These scenarios involve different capabilities, actores, time horizons and interventions. When an argument begins with current language/bio models but eventually depends on a future economy operated by autonomous AI systems, that shift needs to be visible.
The emerging empirical evidence support a more complicated story than either side usually tells. In this preregistered trial involving 153 novices, participants with access to mid-2025 frontier models were no more likely than internet-only participants to complete the study's core lab workflow: completion rates were 5.2% and 6.6%, respectively. However, it is worthy of note that AI-assisted participants made more progress on some intermediate tasks and performed better on the cell culture component, suggesting a modest and task-dependent benefit even though the primary outcome showed no significant uplift.
Across the wider evidence base, studies of computer-based knowledge and planning tasks have found effects ranging from minimal to substantial, while two publicly reported wet-lab studies found no statistically significant uplift on their primary outcomes. Most studies have also focused on novices, leaving the effects on skilled and partically skilled actors comparatively underexplored.
I find it difficult to look at this evidence, albeit dated --as they tasted less capable models compared to Fable 5.1 and Astra -- and conclude that current models have already dissolved the barriers to consequential biological work. I find it equally difficult to conclude that AI makes little difference. The more defensible interpretation is that its effects vary across tasks and actors, with stronger evidence for assistance in computational, informational and planning activities, and much weaker evidence that current systems allow novices to complete complex physical workflows.
The harder interpretative problem is that an unchanged end-to-end success rate does not tell us whether AI changed nothing, or whether it improved one part of the workflow while another constraint continued to determine the final outcome.
In the thread on X, the OPs strongest arguments apply to a particularly demanding scenario: an AI system autonomously operating an integrated virology lab. Showing that this scenario remains technically and economically remote would matter, but it would not establish his broader conclusion that AI has little effect on bio-risk.
The nearer-term question is whether AI can materially improve the performance of actors who already possess some combination of expertise, infrastructure, access and human collaborators. Such assistance could matter without removing the physical bottlenecks, particularly if it changes a consequential decision, reduces failures at a difficult stage, or makes several distributed capabilities easier to combine.
I believe the real conversation should concern how much risk can change before full autonomy, and whether existing safeguards remain effective as the division of work between humans, AI agents and external providers evolve.
This is where my own view has shifted as I have worked through eight AI-enabled biotechnology pathways, tracing how models, biological data, specialist tools, synthesis providers, labs, human decisions and experimental feedback combine to produce an outcome. I kept returning to the same basic point: a model's capabilities tell us surprisingly little until we know who is using it, what complementary resources they possess and which parts of the pathway they can actually complete.
The same model might provide little practical value to a novice who cannot recognize an incorrect answer, yet be genuinely useful to a scientist who has the expertise and infrastructure to act on a narrow improvement in design, analysis or troubleshooting. The relevant question is therefore actor-relative: how does access to AI change what this particular person or organization can accomplish?
AI may also redistribute constraints rather than remove them. Easier computational design could make physical validation the dominant bottleneck; better lab automation could shift the constraint toward access, materials or authorization.
This is why both "the model can do it" and "biology is too hard" strike as me inadequate descriptions. Each captures part of the problem while hiding the rest.
I would like to see fewer general claims about whether "AI can create a biological weapon" and more claims with a visible actor, pathway, time horizon. Before saying that a model substantially increases bio-risk, we should be able to answer five questions:
1. who receives the alleged uplift?
2. what are they trying to accomplish?
3. which part of the pathway does AI change?
4. which important constraints remain?
5. what evidence would cause us to revise our assessment?
We should also say whether our claim comes from observed performance, an extrapolation from current evidence, or a precautionary scenario. There is nothing wrong with discussing scenarios that depend on capabilities that do not exist yet, but the reader should not have to work out whether a claim concerns current models assisting novices, future systems assisting experts, or autonomous agents operating laboratories in a radically different technological environment.
This would make some claims narrower and less dramatic, but it would also make them more useful. Instead of debating whether AI changes everything or biology is simply too hard, we could ask whether AI improves sequence analysis for skilled users, reduces failures during experimental planning, helps coordinate distributed services or weakens a specific safeguard at a consequential handoff. Those questions can be investigated, updated and connected to particular interventions.
Many responses to the discussion on X provide useful examples of how overconfidence can run in both directions. I think many skeptics are right that evolution belongs in the analysis of possible bio-risk from AI, but I am less confident that it points in a reliably reassuring direction. Evolution does not have a general preference for making pathogens harmless; what happens depends on how virulence affects transmission, when transmission occurs, what an engineered feature costs the virus, and the environment through which it spreads. Inserted sequences can be unstable, sometimes very unstable, but there is no general rule that a harmful feature will quickly disappear. The evolution of virulence and the stability of inserted viral sequences are both more complicated than that.
This is another place where everyone needs to show more of their reasoning. risk advocates should explain how engineered properties could survive long enough to produce the projected harm, while skeptics should explain why evolution would reliably remove them. Claims that the benefits of AI-enabled biology will greatly exceed the harms, or that even an enormously destructive pandemic could not threaten civilization, also require arguments of their own; expertise in AI and virology does not settle those broader questions.
The EA community's willingness to take low-probability, high-consequence risks seriously is a real strength, but it depends on being equally willing to separate demonstrated capabilities from extrapolation, and to take seriously evidence that cuts against our expectations.
Telling skeptical scientists that they are focused on today while we are thinking about tomorrow only gets us so far. Any forecast still needs a plausible account of how we move from here to there: which constraints change, which remain, and what evidence would move us. Otherwise, "future capabilities" can become a way of keeping a claim beyond the reach of evidence.
This will continue to matter because credibility affects whether scientists participate in evals, companies adopt safeguards, and governments receive advice grounded in how biological work actually happens. A better debate would be more specific: skeptics should identify the step they think remains prohibitive, risk advocates should explain why they expect it to change, and both should say what would change their minds. That seems like firmer ground on which to take the risk seriously.
I am one of the authors of the aforementioned novice uplift randomized controlled trial (Hong et al. 2025), and I want to clarify that while the study did not detect a statistically significant difference on the full model reverse-genetics workflow, nonetheless we did find uplift on component tasks such as mammalian cell culture!
Here, a reasonable takeaway is that tacit knowledge in biology is not a strong barrier for uplift — instead, participants in the AI arm performed much better with AI assistance, even on tasks that we would consider to be tacit-knowledge heavy.
An important caveat is that the study was conducted mid-2025, using the then-frontier models of that generation but more importantly, with a participant pool that is representative of the average level of AI proficiency of that period. As time goes on, secular trends in AI skill development will likely make people better at using models.
Finally — I'm curious to hear what's the second RCT that you mentioned in your post? I'd be grateful to receive the link to it!