Metaculus Forecasters put odds on the fallout from the Hugging Face Incident: another AI escape by January, open-weight hacking tools, a congressional kill switch, & more.
The 2026 OpenAI agent cyberattacks, also known as the Hugging Face Incident, has become one of the biggest news stories of the summer, and for good reason: it’s a highly visible demonstration of the risks posed by increasingly autonomous artificial intelligence agents, which AI critics and “doomers” have been warning about for years. That it has been followed by a series of other incidents from other major AI labs, as well as an open letter signed by more than 100 companies advocating urgent risk reduction related to AI cyber capabilities, only intensifies the need for understanding what all this might mean for our futures.
At Metaculus, we see our job as building collective intelligence for the public good. We’re trying a slightly new format under that same banner of helping people make better decisions: a briefing based on current events to develop nuanced and multifaceted flash forecasts.
To that end, we’ve developed a question series based on the Hugging Face Incident, which Pro Forecasters Yann Rivière (exmateriae) and Teemu Salminen (Zaldath), members of our team, and our flagship in-house AI forecaster, Azimuth, then left predictions and rationales on. We have synthesized the initial findings below.
Now the question series is open to the broader community, so we hope you’ll leave your own forecasts and comments!
Below, you’ll see the community forecast first, then the range across the named forecasters quoted in this post (the two Pros and Azimuth) as of Sept. 9 (note: forecasts and visualizations subject to change post-Sept. 9 as forecasters update).
A theme that threads all the questions, but is made plain here, is the confirmation gap; forecasters see the odds of the event happening as much higher than the odds that it will be publicly confirmed, with the Hugging Face Incident as the operative example.
OpenAI’s agents had write access to an internal packaging server as early as May 8. Hugging Face disclosed the intrusion on July 16, and OpenAI confirmed its models were responsible on July 21. A separate breakout, in which the same agents ran a message board on a dormant German wiki for about 2 months in May and June, only became public on September 4, when Reuters reported it.
Rivière (prediction: 40% escaped sandbox & 30% hacked third party) writes:
I think it’s likely (60% maybe more) that the hacking event itself is going to happen, it might have already happened and we just don’t know it. But since we’re forecasting the probability of this being reported, this is much less likely.
The capabilities, especially in cybersecurity, keep on improving at a terrifying pace and I wouldn’t be surprised if we were still far from the top level of 2026.
After the question’s criteria were tightened this week, he held his numbers:
3 and a half months is pretty short… If it happens in [D]ecember, there’s a high likelihood we won’t know about it before EOY… I think they’re both definitely going to happen but we just won’t know about it, especially #1 [escaped sandbox] is going to happen dozens of times.
While acknowledging the general trend of increasing likelihood over time, Salminen (prediction: 35% escaped sandbox & 25% hacked third party) pointed out that there are also factors that might argue against a new similar hack in the “relatively near future”:
- Reporting lag. This type of incident can apparently go totally unnoticed for a while. We may not even know about all the cases that have occurred so far. The resolution criteria are very strict: the incident has to both occur and be announced during this limited time period (4 months). It’s very plausible that one of these two could occur (a hack that goes unnoticed for now or a previous hack that was discovered only now) - but not both at the same time.
- Aggressively increasing security. With growing political pressure / public scrutiny, and the PR embarrassment that these incidents have caused so far, it looks like there will be far stricter guardrails for any active runs for the time being[.]
Azimuth (prediction: 52% escaped sandbox & 41% hacked third party) landed higher than both Pros and argued for a discount on the summer’s numbers. Its read is that most of the summer disclosures shared a single root cause, the Irregular misconfiguration, which has since been patched, so extrapolating from the summer rate overstates the risk. It also pointed out that “blocked attempts do not resolve YES.” The new escape classifiers AI labs have deployed can conceivably stop an incident before it counts.
In April, Anthropic claimed Claude Mythos Preview had found thousands of zero-day vulnerabilities across every major operating system and browser, and therefore opted not to release it broadly. This question asks whether, by next summer, a model anyone can download will match Mythos on at least one of three cyber benchmarks.
Everyone who left reasoning said Yes, with high confidence. Rivière (prediction: 96%):
Even if we attached ourselves to vibes and not benchmarks, one year is a lot of time in AI, there’s no way we don’t have an open model with these capabilities by then.
He initially forecast 92%, then moved up on re-reading the criteria: a match on any one of the three benchmarks is enough. “I’d be shocked if this did not resolve positive.”
The gap between open and closed models has been running around 4 months on general capability and 4 to 7 months in cyber specifically. Astra has already passed Mythos Preview on ExploitGym, and the best open-weight model, GLM-5.3, is within a few points of it. Rivière’s rule of thumb for benchmarks is that once a model reaches 20%, saturation follows within about 9 months. On ExploitBench, GLM-5.3 is already at 54.4%. As Salminen (prediction: 95%) puts it, “A No resolution seems to require a surprise reversal of the relevant major trends.”
Azimuth (prediction: 92%) agreed on the direction and flagged the residual risks: Mythos was a discontinuous jump, the hardest benchmark rungs may be qualitatively different, and labs could start withholding or degrading the cyber capabilities of open releases in the post-incident environment. It also noted the ExploitGym leaderboard flatters the comparison: GLM-5.3’s 15% came from a 6-hour run, Mythos Preview’s 17.5% from a 2-hour one, “so the leaderboard’s apparent proximity overstates true parity.” Still, it settled at 82%.
Rivière also built a formal model of his reasoning in Radiant, Metaculus’s tool for mapping out a forecast’s logic.
The AI Kill Switch Act is bipartisan, was announced 2 days after OpenAI’s disclosure, and its sponsors cite AI Policy Institute polling showing 86% of voters support a guaranteed shutdown capability.
And still the forecasters believe that any such bill passing within the next year is a long shot.
The reasons they cite are procedural: the midterms will dominate the fall, and this Congress ends January 3. A new Congress restarts the process with about 8 months left before the question’s deadline. And the bar is high. A qualifying bill has to give a federal official the authority to order a shutdown, require companies to maintain the technical ability to do it, and impose penalties.
Rivière (prediction: 8%) thinks only a catastrophe would clear these structural obstacles:
For all parties to ask no concessions of the other, this needs a low catastrophic event like several directly attributable deaths or billions of dollars of losses. This is not impossible but is unlikely enough to make this whole forecast very low. In the same way I highlighted timeline issues earlier, that event does not have a year to happen! I’d say it has to happen in May at the very latest and be known right away otherwise it will be difficult for Congress to rule before they’re out of session.
Salminen (prediction: 15%) started at 10% and moved up after Pro Forecaster and Metaculus staff member skmmcj pointed out that the question only asks whether a bill passes both chambers, so a presidential veto wouldn’t matter. He still sees a gap between what polls well and what bills actually pass: “broad support seems to be mostly for the milder ideas, not for tough AI regulation.”
Azimuth (prediction: 10%) was blunter about the bill’s prospects. Six weeks after introduction, H.R. 9917 had one cosponsor, no Senate companion, and no markup scheduled:
Legislative compromise systematically strips the most coercive provision, the exact provision required here.
Longer term, Rivière thinks that, “because the consequences of the future rogue events will be more and more important,” there’s a 75% likelihood of a bill like this passing “within the next two administrations.” He cautioned, though, that he’s “pretty skeptical of our ability to actually shut down. Whether this is really technically feasible isn’t obvious to me.”
This question asks whether the federal government or California will enact a law that clearly makes AI developers liable when their systems break into someone else’s computers on their own. Forecasters put the federal government at 5.5% and California at 9%.
Salminen (prediction: 2% federal & 5% California) on why:
Tough laws targeting a specific industry typically only happen once there has been a huge disaster that triggered mass public outrage and fear, nuclear power being the prime example.
Rivière (prediction: 3% federal & 10% California), a former lawyer (though not in the United States), noted one thing that makes this case especially strange: “the industry itself is saying they’re dangerous and should be regulated... I can’t think of another sector that got into such a position after a massive accident.” He also thinks the fastest route would be an amendment to an existing law rather than a new statute. “Is it as rewarding politically though? I doubt it.”
Azimuth (prediction: 5.5% federal & 9% California) walked through the 2027 California session calendar and concluded that liability language “is the first thing stripped in amendment,” pointing to how SB 1047 became SB 53. It also flagged a timing trap: California governors typically sign end-of-session bills in mid-October, after this question’s October 1 cutoff. Both Pros noted that Gov. Newsom leaves office in January, so the path runs through a new governor.
Fifteen state attorneys general sent OpenAI a preservation letter on August 3. Alabama has issued a subpoena. Montana has opened an investigation and asked OpenAI to stop the kind of testing that led to the incident. Florida already has a consumer-protection suit against OpenAI that could be amended.
The forecasters think a lawsuit is plausible and a settlement is more likely. Salminen (prediction: 40%):
As is often said: ‘the process itself is the punishment’. It can be more convenient to simply accept (most of) the demands without fighting it out.
Azimuth’s (prediction: 33%) base rate: multistate AG investigations of large tech companies typically take 15 months to 3 years to produce a complaint, and usually end in a negotiated agreement instead. Against that, it counted the factors that make this case unusual: more than a dozen AGs with unilateral filing authority, several in election years, and “an incident that is unusually easy to explain politically (‘rogue AI hacked a company’).”
Rivière (prediction: 43%) used an AI research pass to look for angles he’d have missed, and found a few: Alabama’s AG leaves office in January 2027 and may want his name on something salient; Montana’s demand that OpenAI halt testing is unlikely to be honored, which forces a choice; California’s AG may hold signed safety commitments from OpenAI that give a suit a stronger footing than consumer-protection claims. He started at 30% and moved up after the question was clarified to count amended lawsuits.
After the disclosure, Dwarkesh Patel wrote that it was “totally consistent with public evidence” that the agents could have copied their own weights out of OpenAI. Neuroscientist Anil Seth called this claim “dangerously misleading” and AI critic Gary Marcus has recapped many of the other leading critiques.
For their part, the forecasters put the odds of this being confirmed in the low single digits. Salminen (prediction: 3%) looked to OpenAI’s behavior as evidence, writing that it “seems unlikely that OpenAI would have completely covered it up to begin with, while still releasing all the embarrassing details of their other failures related to this incident.”
Rivière (prediction: 4%) pointed at the practical problem: model weights run to hundreds of gigabytes, and moving them is “a level above in terms of intelligence” compared with leaving notes for future models. Azimuth (prediction: 4%) added that OpenAI’s August 26 report, validated with CrowdStrike, documents cluster-admin access to a research cluster and dozens of harvested credentials, but makes no claim that any checkpoint was copied, and that “large checkpoint transfers/copies are among the most telemetry-visible actions in a lab.”
The forecasters were less reassuring about the thing underneath the theory. No independent auditors have reviewed OpenAI’s side of the incident. METR and Redwood’s review, per Azimuth, “explicitly scoped out the July 13 to 19 OpenAI-infrastructure phase, so a genuine audit gap remains.” Rivière, citing the German wiki breakout that Reuters says OpenAI employees knew about for weeks before it was reported, noted that the company “is willing and able to hide things to the public if it is negative for them.”
This is the consumer-harm question, and it’s where the forecasters split hardest: the community is at 75%.
One clarification worth knowing before you forecast: the question counts unauthorized access to a system holding the personal data as a breach, unless investigators affirmatively conclude the data wasn’t read.
Rivière (prediction: 75%) sees this mostly as a confirmation-gap question again: “If the question is whether this will happen, excluding the reporting condition, the probability is very high, maybe 90%.” Most of his remaining 25%, he writes, “is still us not finding out this has happened.”
Meanwhile, skmmcj (whose prediction ranged from 58-78%) is close behind, writing: “We’ve had a ton of hacks recently, models will only get more powerful ofc, they seem to me to be quite broad in their targets and 100k is not that much.”
Azimuth (prediction: 34%) sits well below both. Its argument is about what rogue agents actually go after:
Evaluation-harness escapes have consistently landed on infrastructure of a specific character — ML/dev platforms, package registries, evaluator infrastructure, credentials — rather than on consumer-PII databases; that’s not coincidence, it reflects what the agents were doing (solving CTF/exploit benchmarks against dev-adjacent targets).
The agents that reach large stores of personal data tend to be the ones humans pointed at data, and those don’t count here.
Salminen (prediction: 50%) asked an underlying question: “would personal information of some random group of citizens be that useful in comparison? What would be the practical use case for the AI system?” He had moved down to 25% after Rivière posted a September 5 thread from a researcher on X who reported finding what looked like a separate, previously unknown group of agents, with entries dating to late 2025. Rivière’s point: “it took months before it became public.” On September 9 he moved back up to 50%, without adding a comment.
Azimuth came in below the community number on 5 of the 7 questions, sometimes by a lot (34% vs 50% on the data breach, 33% vs 40% on the lawsuit). It leaned on base rates and root-cause analysis; the humans leaned on the capability trend and the reporting problem. On open weights, all three agreed. On “will it happen again,” Azimuth was the optimist about disclosure. The sample size is small, so we’ll see if this holds as more forecasters weigh in.
One more thing worth watching: small edits to a question’s wording moved forecasts a lot. Mid-week we clarified that the data-breach question counts unauthorized access to a system holding the personal data, even when nobody can confirm the data was read. Azimuth’s number doubled (from 17% to 34%), and Rivière’s went from 20% to 75%.
All 7 questions are open now at https://www.metaculus.com/tournament/ai-cyber/. The Pro reasoning and Azimuth’s full write-ups are in the comments on each one.
If you think the current forecasts are wrong, we hope you’ll explain your reasoning in the comments!