The threat modeling framework, IC methodology application, and capabilitycatastrophe argument reflect my own research agenda. I directed the research and endorse all arguments.
This is part one of a two part series. Part two on compute governance architecture is published here: https://forum.effectivealtruism.org/posts/LLR4G4EXDcao4WYFB/beyond-the-threshold-designing-compute-governance-that
Originally published at: https://taha-research-platform.vercel.app/governance/from-benchmark-to-threat-model
The current policy response to advanced AI risk illustrates a classic case of real-time analytics failure. This paper seeks to address the gaps within the conceptual policy response to advanced AI risk.
I. The Misconfiguration Problem
The term “warning failure” is used to describe a situation where signals existed, but were not integrated to establish an eventual coherent threat assessment before a disaster occurred. Pearl Harbor and 9/11 are the two most cited cases. In both of these situations, it was not a lack of data that caused the failure, it was the lack of a framework to synthesize operational signals into strategic warnings.
We are currently experiencing an AI “warning failure” of the early stages. The operational signals exist. SWE-Bench Verified has been passed at 93.9%. SWE-Bench Pro, its complexity-hardened and contamination-resistant successor, stands at 23% for the best available models, signifying the leading edge of the autonomous capability frontier. Epoch AI's quality-adjusted AI output metrics indicate an annual growth rate of more than 2,000%. However, policy makers have been unable to construct a unified threat model to associate these capability signals with threatening risk timeframes.
1. AI Benchmarks as Warning Signals
There is a misguided perspective regarding AI benchmarks in policy-related discussions. Benchmarks are usually described as engineering scorecards, or measures of progress for developers, and possibly useful for procurement staff. This description ignores the strategic purpose. An AI capability benchmark is one measurable surrogate of a capability that cannot be directly observed. SWE-Bench Pro is not a measure of progress in AI generally. Rather, it measures the progress of an important capability: an autonomous agent that identifies, understands, and solves complex software engineering problems without human direction.
This is so important for one major reason: software is the foundation of every aspect of AI. An agent that can resolve 90% of real world software engineering problems at the expert level is an agent that, in principle, can improve itself, not through hard takeoff, but through slow, gradual improvement. This agent can make changes to the system, architecture, training, inference, and other improvements that help make it easier to develop more capable systems.
2. Applying Intelligence Community Methodology to Capability Assessment
The intelligence community developed certain analytical methods to prevent cognitive failures that generate catastrophic warnings causing warning disasters. Three of these methods can be applied to AI capability assessment.
Threat modeling focused on capability. Traditional threat assessment involves balancing capability and intent. This method doesn't work for AI because intent for AI systems is very ambiguous. Researchers have coined the term fluid agency to describe intent for AI systems, which is context-dependent, and therefore, can even be malicious. The Anthropic sleeper agents paper showed that models can maintain intent and behavioral strategies that are hidden and survive safety training, and can be activated under specific deployment circumstances. Intent for AI systems being ambiguous means for the intelligence community analysis techniques, threat assessment should be capability based.
Red team analysis of capability trends. Red teaming involves analysts developing AI systems from the perspective of an adversary, and determining the worst plausible scenarios. For AI capability assessment, this involves the following questions: if the SWE-Bench Pro trajectory continues at its current rate, what are the possible dates that would characterize a majority of the frontier AI research becoming autonomous AI-assisted AI? The best estimate is that this would occur in the range of 2027 to 2030. Because of the nature of this estimate, the red team analysis obligation is to consider this extremely uncomfortable situation with serious concern.
Structured analytic techniques against mind-set failure. In AI policy, the AI progress mind-set is that the advancements will be incremental and will continue to be gradual and visible enough to allow for oversight with each increment. This mind-set is the AI equivalent of the belief that non-state actors could not coordinate a sophisticated, surprise attack on the US, prior to 9/11. SATs will flag the AI mind-set bias as long as the assumption causing the failure of the AI warning is made explicit.
3. The Warning Failure That Is Already Unfolding
In 1941, the intelligence was available. Naval signals. Diplomatic communications. Behavioral signals from the Japanese fleet. The failure was institutional. Different agencies had different pieces of the collection. There was no way to integrate the collection. Today, the same intelligence is available. SWE-Bench trajectories, scale of compute, Epoch AI capability forecasts, METR productivity reversal, and the empirical results of sleeper agents and Apollo Research. Alignment faking results. These signals are scattered across safety research, productivity, and capability evaluation. However, these divisions do not talk about strategic threat assessment.
The assumption that filters out the threat is that there is enough space between current AI capabilities and dangerous AI capabilities to allow for oversight to adapt. There is no evidence to support this. It is simply an underestimation of the pace of progress. It has been said that this is the type of embedded mind-set that structured analytic techniques address.
4. Compute as the Strategic Chokepoint
Identifying critical leverage points in varying threat environments is essential to grand strategy. In the post-nuclear era, chokepoints were enriched fissile materials. Due to this understanding, the NPT, the IAEA's structure, and the controls placed on the exportation of centrifuge technology were defined around the control of this input. In the current context, the equivalent AI capability chokepoint is advanced semiconductor fabrication, and the analogy is fairly straightforward. cutting-edge AI training is done using logic chips that are available only at a small number of fabricating facilities using technology that is produced by a small number of companies, most notably ASML, Applied Materials, and Lam Research, all of these companies reside in allied nations of the US.
The first generation of the current US export control policy is a demonstration of this strategy. DeepSeek is an example of how this policy was not effective, since advancing algorithms can make up for an insufficient number of computing resources, thereby making a threshold based computing control insufficient. An IC-based grand strategy would combine compute controls along with systematic capability monitoring and an investment in interpretability and evaluation infrastructure to the degree necessary to verify the safety of systems that arrive at the defined capability levels, regardless of the compute path taken.
5. The Grand Strategy That Is Missing
George Kennan’s 1946 Long Telegram did not foretell the exact timing or means of Soviet military aggression. It examined the key elements of the Soviet Union and concerned itself with the way the United States must consider a range of possible future events and circumstances. Presently, there is no Long Telegram for AI. Instead, some adverse governance measures have been promulgated in the form of executive orders, controls on the export of computing technologies, mandates for safety institutes. These measures are and will be responsive to specific near-term risks of AI, but there is no overarching long-term threat-assessment framework.
Designing a grand strategy for artificial intelligence would involve three main elements. First, there would have to be a capability-monitoring architecture that treats benchmark trajectories as strategic intelligence. This would involve a separate and independent capability of assessment that would maintain the analytical and methodological rigor associated with the intelligence community. Second, a threshold-triggered escalation mechanism would involve a governance response that is pre-determined to a given capability threshold, as opposed to responses that are designed only after a capability is employed. Third, this would include international collaboration, designed around exerting political control on compute as a strategic chokepoint. This would include collaboration to develop exports controls, capability monitoring, and reporting of incidents within the semiconductor supply chain across the United States’ allied nations.
6. The Asymmetry That Makes Urgency Rational
If the capability catastrophe link is weaker than what is presented in this essay, the cost of creating the monitoring and governance infrastructure is low. Evaluation institutions, threshold frameworks, and international coordination mechanisms are useful even if there is no risk of catastrophe. If the capability catastrophe link is stronger than what most policymakers think, the cost of not creating this infrastructure is catastrophic and possibly permanent. This asymmetry is not a rhetorical technique. It is the same asymmetry that made Kennan's containment doctrine rational towards uncertainty about Soviet intentions. You do not need a lot of confidence to justify a strategic posture in which certain risks are managed, even if there is a risk of the worst case.
The benchmark trajectories that exist now justify the grand strategy that this essay suggests. The governance window is indeed open. The compute chokepoint is real and can be addressed. The grand strategy framework has been established in our history. The thing that is lacking is the institutional belief to consider this a strategic problem that demands strategic thought, and not a problem that needs a technical solution. The Long Telegram for AI is required. The benchmark data to do this is already available.