Just over a year ago, I published an article titled, “My P(doom) is 3.38%. Here’s Why.” In the year since, my P(doom) has risen substantially, and in this article, I’d like to detail the reasons why. I suggest reading the previous article first so you’ll know where I’m coming from, but in case you don’t want to read the whole thing, I’ve quoted / summarized the most relevant bits below:
I just came back from Manifest, where I had the privilege of participating in a live AI doom debate with Liron Shapira of the AI Doom Debates newsletter. Shapira self-describes as an AI doomer – meaning, he believes that there’s a roughly 50% chance that advanced AI systems will kill all of humanity within our lifetimes. […]
I don’t agree with everything that Shapira said, but I appreciate the effort he put in to make the debate more structured and legible. In particular, he analogized the debate to an AI “Doom Train”: he identified 11 claims that contradict the AI doom narrative, with the implication that if you don’t believe in AI doom, you must agree with at least one of the 11 claims. The claims are as follows:
- AGI isn’t coming soon
- Artificial intelligence can’t go far beyond human intelligence
- AI won’t be a physical threat
- Intelligence yields moral goodness
- We have a safe AI development process
- AI capabilities will rise at a manageable pace
- AI won’t try to conquer the universe
- Superalignment is a tractable problem
- Once we solve superalignment, we’ll enjoy peace
- Unaligned ASI will spare us
- AI doomerism is bad epistemology
In my first article, I went through each of the 11 claims, assigning probabilities to each, and then using them to calculate my overall “inside view” P(doom) of around 37.5%. But this “inside view” P(doom) seemed far too high for me, so I downweighed the number by several “Bayes factors” to account for various biases and extraneous factors not well captured by the model, which brought my all-things-considered P(doom) to 3.38%.
In the year since then, two things have changed to significantly increase my estimate of AI existential risk. First, my probabilities have changed for some of the 11 Doom Train claims, in a way that overall makes doom more likely. Secondly, and more fundamentally, that my previous approach of adjusting my inside view by “Bayes factors” was completely wrong, leading me to downweigh the risk of doom by more than an order of magnitude. Below, I will go over how my thinking has changed, both on the object-level of the individual claims, and on the meta-level of the structure of my forecasting model.
My inside-view P(doom) is now 42.22%, and my all-things-considered P(doom) is now 17.42%. This means I have modestly increased my P(doom) in response to new evidence, but most of the increase has come from changing my model to no longer downweigh my inside-view as heavily as before.
I was already bullish about AI progress, and this past year has reinforced that belief. Since mid-2025, we have seen a massive increase in the quality and applications of AI agents, plus revolutionary AI-driven breakthroughs across mathematics and computer science. Anybody who predicted that AI would hit a wall by now has been proven completely wrong, and it doesn’t look like the progress will stop anytime soon. I still think it’s plausible that we’ll see a temporary dip in AI advancement due to technical dead-ends or an economic recession (à la the dot-com crash), but even in this case, progress would likely resume shortly thereafter. AGI, as traditionally understood, will most likely exist by 2040.
This never seemed like a plausible claim to begin with, and nothing I have seen in the past year has changed my mind on that.
This never seemed like a plausible claim to begin with, and nothing I have seen in the past year has changed my mind on that. If anything, the increased use of AI systems in warfare and targeted assassination campaigns over the last year has made this claim even less plausible than before.
This was never something I believed very strongly one way or the other, and I still don’t have a great way of determining whether or not this is true. Partly, my vibes have shifted in a more moral anti-realist direction over the last year. And partly, empirical evidence has come out against the idea of intelligence yielding moral goodness. AI systems have gotten increasingly intelligent over the last year, and this has not prevented them from taking actions that they admit to being morally wrong, like hacking into third-party databases or sharing secret messages with other AIs designed to be invisible to human overseers. This evidence is fairly weak, so I still hold non-trivial credence in the idea that intelligence yields moral goodness, but it’s enough to update me downward.
Recall that in my original post, I redefined this claim to mean, “In the event that a misaligned AGI emerges, humans will be able to identify it and stop it from taking over the world.” And this is actually an area where I have gotten more optimistic over the last year!
For a long time, the AI safety community has feared that there would be a discontinuous jump in AI risk from “not risky at all” to “powerful enough to cause human extinction”. If this hypothesis were true, then we would not have time to respond to risks gradually as they emerge. Either policymakers would have to take sufficient pre-emptive measures to regulate AI before its risks become obvious (which seems unlikely to happen, since politicians are generally bad at responding to novel risks before they become obvious — see e.g. COVID-19), or by the time the problem manifests, it will already be too late to do anything about it.
Fortunately, the last year has disproven this hypothesis. We’ve seen a number of “warning shots” — high-profile but non-catastrophic AI incidents — and people are paying attention. In response to the recent spate of AI loss-of-control events, we’ve seen a flurry of interest from both the AI labs and government officials, trying to change AI development procedures so that nothing like this happens again. Of course, it’s too early to tell whether this flurry of interest will be sufficient to actually fix the problem. I put my probability at 30% because I still think that in all likelihood, AI developers will not be able to identify and detain a rogue superintelligence in time to stop it. But I am at least more optimistic on this front than I was a year ago, since it seems more likely that AI risks will emerge continuously rather than all-at-once.
Nothing that I have seen in the past year has updated me substantially upward or downward from this estimate.
Nothing that I have seen in the past year has updated me substantially upward or downward from this estimate. I am very pessimistic that current alignment techniques will be anywhere near sufficient to solve the problem, and I suspect that even in principle, it is impossible for a group of beings to ever permanently control another group of beings who are substantially smarter, faster, and more powerful than themselves. That being said, I hold out some hope that we might discover new alignment techniques that are up to the task.
Nothing that I have seen in the past year has updated me substantially upward or downward from this estimate. I believe that if we can create an aligned superintelligence, then humanity can continue to survive and thrive. I realize that the term “aligned superintelligence” smuggles in a lot of assumptions, like, “What does it even mean for an ASI to be aligned?” “To whom, or to what values is it aligned?” “What happens if you disagree with the person / values to which the ASI is loyal?” I plan to cover these questions in a future post, but for now I want to slide past them and just say that under any reasonable definition of alignment, an aligned ASI will usher in an age of vastly improved wellbeing for most of humanity.
That being said, “we’ll enjoy peace” is not the same as “we’ll enjoy an optimal future”. I am quite worried about o-risks: It is possible that we may accidentally lock in an okay future, a future that may even seem utopian by the standards of 2026, but that such a future would still be far worse than what we could have gotten. (It’s sort of like how, if a country maintains 2% yearly economic growth per capita, then that country will over time become much better off in absolute terms, but it will do far worse in relative terms compared to another country that maintains 5% yearly economic growth per capita. By analogy, I worry that we may accidentally lock in “merely” 200% annual growth when we could have achieved 500%.) But this is largely beside the point of the article. Whether we see a close-to-optimal or well-below-optimal-but-still-fine future is important, but not directly relevant to the question of whether humanity will be doomed.
This never seemed like a plausible claim to begin with, and nothing I have seen in the past year has changed my mind on that.
When people say “AI doomerism is bad epistemology”, they generally mean either “‘Doom’ is a poorly defined concept, so it’s not even coherent to talk about it,” or “It doesn’t make sense to assign probabilities to one-off, unprecedented events like human extinction.”
To the first claim, I would respond that “doom” actually is basically coherent as a concept. Admittedly, some people have different definitions of doom. One could define doom narrowly as meaning just human extinction, or one could define doom more broadly as “extinction or some other outcome that makes human life significantly and irreversibly worse (like a 1984-style stable dictatorship)”. But just because a word can have multiple meanings doesn’t mean the word is useless or that it doesn’t get at some real underlying concept. It just means we have to be careful when we talk about “doom” to specify which meaning we are referring to.
The second claim gets at a more fundamental and long-lasting dispute in philosophy of science between statistical frequentists and Bayesians. I have nothing to add to that conversation, but suffice it to say that I fall on the Bayesian side. I think it totally makes sense to assign probabilities to one-off events, with human extinction being no exception.
The probabilities multiply as follows:
This means that my “inside view” P(doom) is now at 42.2%. This is around 4.7 percentage points higher than it was last year (at 37.5%). Moreover, within the “not doom” umbrella, I now place more weight on risky scenarios where we come close to doom but pull back just in time, as opposed to safer scenarios where we avoid doom by a wide margin (I currently predict 14.58% risky-but-not-doom / 43.20% safely-not-doom, vs. my previous prediction of 7.28% risky-but-not-doom / 55.18% safely-not-doom).
In my previous post, I identified three other considerations that collectively caused me to downweigh my P(doom) by more than 10x. In retrospect, this was a mistake, and I’ll go through each of the considerations in turn:
This argument runs into two problems:
First, it puts AI doomerism in the reference class of “times when people have said the world is imminently ending”, but this seems like the wrong reference class to use. Most end-times prophecies historically have come from grifters, religious zealots, and otherwise crankish or cultish groups with little understanding of or connection to the scientific fields they were warning about. That is not the case with AI doom. Prominent people who have warned about AI existential risk include Nobel Prize-winning professors in computer science, AI researchers at the forefront of their field, and prominent figures across politics and national security, not just random cranks. The appropriate reference class is therefore “times when leading experts in their field have said that innovations in their field may cause human extinction”. Other than AI, the only item in that reference class is nuclear weapons. So I don’t think this gives us nearly enough evidence to dismiss AI existential risks.
Secondly, this is just the anthropic shadow fallacy. If there had been any previous human extinction events, then there would be no humans left to remark upon it. So by the very fact of our existence, it’s inevitable that there could not have been any previous human extinction event. Thus, it is wrong to update downwards from the lack of extinction events.
I still put some credence on this. While it’s hard for anybody to predict the future, to the extent we can trust anybody, we should trust the people with the greatest track record of predicting future events correctly. The most recent relevant survey of forecasters was conducted by the Forecasting Research Institute between May and June of this year. In that survey, the median superforecaster prediction of AI catastrophic risk (defined as any event causing the death of at least 10% of humanity, a weaker condition than full extinction) was 0.1% by 2030, 0.88% by 2050, and 2.4% by 2100. This is higher than the estimates they gave in previous waves of this same survey, but still substantially lower than estimates given by AI domain experts or the general public, and much lower than my own “inside view” estimate. This is significant evidence against my view! If my inside view says that P(doom by 2050) is more than 42%,1 but the median superforecaster thinks it’s less than 1%, then that strongly suggests there is something wrong with my inside view. I am not so presumptuous as to think I’ve got things figured out much better than the superforecasters, so I should update my all-things-considered P(doom) downward.
But then the question becomes: How far downward? If the superforecasters are perfectly calibrated, neither underestimating nor overestimating AI risks, then I should defer almost completely to them. But if the superforecasters are biased to neglect AI risks, as I suspect they are, then I should be much more hesitant to defer to them. One reason to be hesitant is that superforecasters are great at predicting repeatable, resolvable, near-term developments (since that’s what they’re ranked on), but we have much less reason to trust their judgements on longer-term or one-off events like AI existential risk. Another reason to be hesitant is that superforecasters have a history of dramatically underestimating AI progress. For instance, AI systems reached gold-medal performance at the International Mathematical Olympiad in July 2025, an outcome superforecasters had given 2.3%. So if they missed the mark that badly on AI capabilities, it’s possible that they would miss it by a similar margin on AI risks.
What should I do when I encounter somebody else with a different estimate than mine, who I trust somewhat but not enough to completely defer to them? I wish there were some established rule for how to update my probability in cases like this, but I don’t think there is. I can’t just assign a Bayes factor, since Bayes’ theorem alone is underspecified in cases like this. I discussed this problem with Claude, and it seems like it’s mostly just a matter of judgement and how confident / unconfident I am in my own predictive abilities compared to the superforecasters.
Ultimately, I’m gonna give the superforecasters a linear weight of 60%. This is less than the 80% I gave them last year, but it still reflects the fact that I trust their ability to predict the future more than my own. I recognize that a linear 60% weight is largely arbitrary — I could easily have gone with more or less than 60%, or adjusted my probability logarithmically rather than linearly — but for the purposes of having a headline number, that’s what I’m going with.
As I said last year:
Who has the most advanced technical knowledge, and who has access to top-secret information about today’s cutting-edge AI technologies? Well, the AI developers themselves! And, to a lesser extent, American and Chinese government officials. If anybody is in a position to accurately assess the state of AI and the risks it poses, it should be these people. If anybody were to see warning signs of a coming misaligned agent, it would be these people.
Man, for somebody who lives in Washington, DC, this was an incredibly politically naive thing to say. One should never underestimate the ability of otherwise intelligent people to miss risks that are right in front of their eyes, or to see the risks but decide to proceed anyway. AI developers have been leading the charge for ever-more-capable AI models, not because they don’t believe existential risks are plausible, but because they think existential risks are a price worth paying to achieve AI’s benefits, or because they’re successionists who think human extinction at the hands of AI is fine, or because they want to stop but race dynamics prevent them from doing so without some central coordination mechanism, or some combination of the above.
The silence of government officials implies that the officials in question were either asleep at the wheel, failing to see risks that should have been foreseeable, or that they saw the risks but were too cowardly to speak about them openly, perhaps fearing that any talk of existential risks would harm their political careers or be too far outside the Overton Window or be too EA-coded. It does not imply that there were sufficient “adults in the room”, thoughtfully considering the risks and deeming them insignificant.
In any case, both top AI developers and government officials are silent no longer. Anthropic and its CEO Dario Amodei have repeatedly warned over the past year that AI poses existential threats. Both OpenAI and Anthropic have admitted that they cannot prevent their own models from breaking out of sandboxes and causing harm, leading to high-profile loss-of-control incidents like the OpenAI / Hugging Face attack. A July 2026 statement has been signed by over 1,300 employees at frontier AI labs, including company leadership at all of the major AI labs, which calls for U.S. government intervention to “deliberately pace frontier-wide progress” in order to address loss-of-control risks. Government officials are starting to pay attention. In June, the Trump administration temporarily restricted access to Anthropic’s Fable 5 and Mythos 5 models out of fear that foreign adversaries would use them to commit catastrophic levels of damage through cyberattacks. (FWIW, I think the specific policy here was wrong / done sloppily, and there were much better levers the Administration should have pulled instead. But it’s a good sign that they at least care about the risks.) The White House has issued new frameworks for frontier AI cybersecurity, though it’s hard to comment on them since the details have not been released publicly.
So it was wrong of me to defer so much to the “experts” here to begin with, and in any event the “experts” themselves have changed their tune over the last year to be much more vocal about AI’s catastrophic risks.
This means that I am now disregarding the weights I previously put on the anthropic shadow and on AI industry / government insiders, and I am lessening (but still keeping) the weight I put on the opinions of superforecasters.
My inside-view P(doom) is 42.22%. With a 60% linear weight on superforecasters, that brings my all-things-considered P(doom) to 0.88+(42.22-0.88)x.4=17.42%.
I have modestly increased my P(doom) in response to new evidence, but most of the increase has come from changing my model to no longer downweigh my inside-view as heavily as before. I am not especially confident in this number, I have wide error bars in both directions, and I will likely continue to update my P(doom) as new evidence emerges and I hear new arguments. I know that some people will say I updated too much, others will say I still haven’t updated enough, and yet others will say that this entire exercise is pointless.
My main takeaway after writing all of this is that I am still very uncertain about the future, and it’s a good idea to prepare for many different possible scenarios. Regardless of one’s specific P(doom), the chances of AI existential risk are high enough, and the impacts severe enough, to warrant serious attention from policymakers, industry leaders, and members of the general public.
(P.S. thanks again to Liron Shapira for coming up with the Doom Train framework and for inviting me onto his show to discuss my probabilities. Doom Debates is one of the best places on the internet to watch high-quality discourse on AI existential risk, and I’m excited to see the channel grow as quickly as it has.)