Dario Amodei's essay We Must Pace the Frontier proposes three steps for achieving it:
I think that the steps are right - but also that the essay stops exactly where the hard details begin. Every step names something to do, but leaves the mechanism blank.
A series of papers I have been publishing since the summer tries to answer these questions. With specific details and models. Here is what each of the steps maps to in these papers:
A pacing agreement among rivals can fall in one of two ways.
First, one party can pull so far ahead that the others might be unable to hold it to the agreement. (The essay's third step fears China in that aspect. However, this applies just as well to the companies in step two.)
Second, the verification behind the agreement might be unable to spot a concealing party,
Any of these will turn the agreement into a mutual suspicion - that is, will effectively void it.
Both failures are underlied by the same: the margin by which the compliant inside the coalition outweigh a member or a bloc that might defect. If this margin is kept above a defended floor, the defection stops being an existential threat. If you lose it, no verification will be able to save the agreement - in the real world, a party that can defect safely will eventually do it.
The paper A Council of Rivals (The Convoy Papers, No. 3; Zenodo 10.5281/zenodo.22019542) describes a design for keeping that margin - including the instruments the essay's steps are missing by now.
The essay asks the frontier companies to agree on "limits on the rate of unchecked AI progress" and sketches "checkpoints": if a model has capability X, it must be accompanied by certifications Y and Z. However, it is not clear what will impose these certifications.
The Council paper proposes a bounded release window - a published maximum lag between any member's frontier and what the other members can access. And calculates for every member an allowance, based on the margin, from data that is public - every member can calculate the allowance for themselves or every other. This allowance is a simple mechanism, which makes the coordination among rivals tractable - unlike with standards like a certification regime, where a judge must decide whether the standard is adhered to.
Capping each rival's holdings at the allowance also caps what any coalition of the defended size can hold. Thus, the coalition does not need to detect a cabal in order to protect against it.
The Council has a defined menu of remedies, which is engaged automatically by a member who drifts above their allowance. The menu includes:
A member is free to calculate its own least-costly combination of these, that decreases the distance between this member and the others. They do not need to release to the Council their reasons to choose exactly this combination, which protects their trade secrets.
When no combination is able to bring a member within a safe distance from the others, a pre-committed escalation ladder is engaged. It runs from mandatory leasing through divestment and field elevation, to a referral to the standing security order outside the Council. These appear unpleasant, but in reality they would be preferable to an overreaction by frightened peers and other parties. That is why a reasonable rival would join the Council.
The essay's antitrust waiver is the legal precondition for exactly this kind of body.
Anthropic commits to this step unilaterally, and rightly: without it, nothing is verifiable. However, an evaluator with employee-level access is only one observation channel - others exists, every one blind to some things, but achieving good transparency together.
The Verification Model (The Convoy Papers, No. 4; Zenodo 10.5281/zenodo.22838800) treats verification as a budget rather than a binary. It envisages nine observation channels:
It also prices them by the fraction of capability a concealer could hold behind each. It catalogues the concealment plays with their counters and residues. And it combines the channels until bringing the residual below what the margin arithmetic tolerates.
The output of the combination is not a binary "verified" or "not verified". It is a numeric evaluation how large a hidden programme could be, and the size of a reaction buffer sufficient for it.
Two of the results of that solve other concerns of the essay. The gap between the margin at which a coalition engages a diverging member and the size of the margin it defends, serve two purposes at once — a concealment allowance and a reaction buffer. This determines how wide that gap must be.
The only attack identified by now, that the Council design cannot solve, is a private efficiency breakthrough, tested below the tracking threshold and applied at scale. It changes a member from being compliant to having decisive advantage between two evaluations. The design bounds that attack to a single evaluation cycle, but cannot prevent it.
The essay's condition for any global agreement — ironclad verifiability or limits tight enough that defection is not existential — is The Verification Model's central idea written as policy. Verifiability is the VM's budget, "limited enough" is its margin, the relation between them is what the VM computes. An agreement can trade one for the other along an openly published curve, and both sides can see where their positions stand on it.
Two further pieces of the Council design speak to this step. There is a pre-committed rule that releases the brake automatically, by indicator, when a non-member closes on the members' frontier. This answers the essay's worry that pacing hands the lead to an unpaced competitor. And the essay's "speed limit" on recursive self-improvement, offered with the SALT treaties as analogy, is named by the Council design "the Continuity Limit". This is the rate of change past which no verification or governance loop can keep up - the natural object of a Level-3 agreement. It is also the subject of a companion paper now in preparation.
Governance proposals usually ask to be believed. The Council proprosal comes with models. The Verification Model's supplements include the normative specification of the reference algorithm, an implementation that passes 2,360 conformance assertions, and an archive of the compute-dynamics models that determined the numbers in the proposal. Including a descriptions of some that failed to achieve the goal:
A reader who doubts a rule can run these models and seek mistakes in them.
The convergence between the essay and the Council proposal invites the provenance question. Information on this:
The ideas that these papers are based on were first published in a forum at kurzweilai.net, in a thread named 'The Singularity Race', during 2015-2016. The forum and this thread have been archived since then, but a snapshot is available at archive.org: https://web.archive.org/web/20180124123624/http://www.kurzweilai.net/forums/topic/the-singularity-race
The essay this post responds to is dated September 2026. The convergence is real and, I think, encouraging: when a problem's shape is clear enough, people arrive at it from different directions.
The essay says the steps "do not need to be taken strictly in order." Neither do these in the Council proposal.
I would welcome collaborators, critics, and anyone who can show where the models break.
The Council papers can be found at https://www.stremsky.org, and at the Zenodo links above.
Further companions — the window's economics, the Commons, the membership, the Council's organs and the endgame — are in preparation.