By Tom Reed | Watch on Youtube | Listen on Spotify | Read transcript
You could probably find a list of less than 10 people in the world where, if you could get them to agree to slow down, you could do it. It’s not some extremely enormous, impersonal sea of people you have to get to coordinate. It’s lab CEOs, potentially people in China, leaders of a couple of countries. … None of the lab CEOs want to hear that argument. They only want to do the non-unilateral things. — Geoffrey Irving |
When should governments slow the race toward superintelligence?
According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.
Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years.
| Want to work with Geoffrey to help align superintelligence? Resolution is hiring! |
The leading AI companies all have broadly similar plans for keeping superintelligence under control:
Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence:
In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted.
This episode was recorded on June 29, 2026.
| Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator/associate, or special projects associate/analyst. |
Our production team includes:
The interview in a nutshellGeoffrey Irving, cofounder and chief scientist of the new research organisation Resolution, expects full-blown superintelligence could arrive within two to three years. He thinks AI companies’ alignment plans might work — but we lack compelling evidence that they will survive the crucial transition from human-level to superhuman systems. His prescription is to slow AI development now, while pursuing a much broader portfolio of theoretical and empirical alignment research. Superintelligence could arrive within two to three yearsThe largest uncertainty is whether models will remain much better at tasks with easily verified answers than at “fuzzy” tasks requiring judgement, intuition, and long-term planning. Geoffrey thinks people put too much weight on this potential bottleneck:
Geoffrey readily allows that progress could instead take 10–20 years, and hopes it will. But he thinks the breadth of current AI R&D and software-engineering capabilities makes a rapid transition disturbingly plausible. AI companies have plausible alignment plans — but little evidence they will scale to superintelligenceThe major companies’ safety strategies combine three elements:
This mixture could work. But all the evidence comes from models that remain subhuman in important respects. Geoffrey expects a significant change when models become better than their human supervisors — precisely where existing experiments stop being informative. He is especially sceptical that good behaviour will automatically generalise:
Capabilities and alignment are also importantly asymmetric: if a capabilities experiment produces a weak model, developers notice and try again. A sufficiently serious alignment failure may be irreversible. AI progress should slow down now, and safety talent should shift to governmentAsked when governments should intervene, Geoffrey’s answer is: now — or preferably sometime in the past. He distinguishes three levels of action:
Coordination may be difficult, but the decisive group could be surprisingly small: fewer than 10 company CEOs and political leaders might be able to substantially slow frontier development. Geoffrey is relatively unconcerned that a pause would halt economic growth. Current models remain far from fully adopted, creating a large “product overhang”: society could spend years learning to use existing systems more effectively even if new model training stopped. He also thinks that alignment researchers in AI companies face diminishing marginal returns and should strategically move to government or independent nonprofits:
Resolution will use theory to “buy the future”Frontier-model experiments can only approximate superintelligence from below. Resolution will complement them with scaled-down experiments and mathematical models designed to capture the distinctive problems that arise above human level. Its initial portfolio covers:
The aim is not necessarily to produce a complete proof of safety. Useful results might instead show that one family of algorithms works while another fails under a particular obstacle — giving companies concrete guidance about what to scale up or abandon. Geoffrey thinks theory is dramatically underexplored because AI companies have repeatedly succeeded through empirical tinkering. But theoretical work may be unusually automatable: models can propose conjectures, search for counterexamples, run numerical experiments, and verify proofs. Resolution therefore plans to combine excellent mathematicians, physicists, and computer scientists with substantial AI automation. Negative results are valuable too. Demonstrating that an alignment approach encounters a fundamental obstacle could justify slowing down, concentrating resources on alternatives, or ruling out an unsafe training procedure. Successful alignment could rapidly transform almost everythingIf superintelligence is aligned, Geoffrey expects extraordinarily fast scientific progress — potentially including solutions to ageing, advanced nanotechnology, bug-free software, space colonisation, and human mind uploads within years or decades. His reasoning is that scientific thought already relies heavily on heuristics rather than explicit step-by-step reasoning. Superintelligent systems could combine far better heuristics with massive automated experimentation and increasingly accurate simulations. But pure market forces would not guarantee a good human future. Once machines can produce and consume everything themselves, humans need no longer be economically important. Preserving human access to resources, political influence, individuality, and meaningful lives would itself have to be part of the alignment target. |
Tom Reed: What’s your rough guess of what OpenAI, Anthropic, and DeepMind’s strategy is for dealing with [superintelligent misalignment]? How do you think they’re going to solve it?
Geoffrey Irving: I think it is all some version of we will do some character training — and they have different approaches there — plus some version of scalable oversight, plus a lot of monitoring. And maybe that monitoring is a mixture of white-box and black-box and so on.
That is:
- Trying to construct environments and training procedures where the models are supervising themselves, so we can kind of keep pace with models as they get stronger.
- Trying to kind of shift the models to be generally good in some way, in such a way that, as they’re supervising themselves, they do that in good ways and that continues.
- And then watch them very closely via AI control and interpretability and so on to again try to catch evidence of bad behaviour and then stamp it out as it is caught.
I think that could work. I don’t think we have a strong argument that the pragmatic mixture of approaches will get all the way there, but it just seems very dicey, and our understanding of the dynamics involved is very weak.
It is interesting that, for example, the different labs have chosen quite different approaches technically to safeguards, they’ve chosen quite different approaches technically to character training. We might need a more rigorous understanding of how those approaches will work if you push them further ahead than the labs can currently see, because all of their evidence is not on superintelligence currently.
Tom Reed: What’s your model of why [AI companies] are more optimistic about it than you? Did you and [Anthropic CEO] Dario [Amodei] already disagree in this exact same way in 2017? Is this something that’s happened in the past few years?
Geoffrey Irving: Turns out we actually did. So Dario, I think from back in OpenAI times, had a take that you train the model on a bunch of good behaviour, and then you scale it up and it will generalise to good behaviour. We literally sketched this on blackboards back in 2018 or 2019. I don’t remember when exactly.
My take is there’s just clearly some notion of phase shift that’s going to happen when you go from human level and pre-human level up to superintelligence. None of the data you have is on that distribution. The question is, will you kind of jump in the right direction or not?
I’m a bit more distrustful of generalisation than I think a lot of the people at labs currently. Some of that is from experience of training models.
Here’s a fun story. In the Sparrow project at DeepMind, we had a model that was fairly good at avoiding saying horrible racist things, but mostly was trained to answer factual questions about the world. This is back in maybe 2022 or something.
Then we said we wanted it to be good at poetry too, so we trained it on some poetry, and then it would do poetry, it would do the questions. On factual questions it would be not racist; it was very happy to write incredibly horrible poetry about racism. You train as best you can on this mixture of abilities, and then you put it in some dramatically new domain, and the generic thing you have to do is then change your algorithms or change the data or something, or it can generalise in kind of horrible ways.
I think there is kind of an intrinsic, maybe evaporative cooling effect of how much do you believe in generalisation going the right way? …
Maybe I think it is the case that as the models get better, they get better generalisation, but we shouldn’t be banking on that to the degree that we are.
Tom Reed: You’ve got this great blog post from several years back where you make an analogy between LBJ’s presidency and aligning superintelligence.
And your point, if I understand it correctly, is LBJ, he’s motivated almost exclusively by power and wanting to acquire more power. He also has all sorts of asymmetric advantages against his opponents, where he’s better at being a politician than them. And yet the American political system still aligns him towards great positive outcomes like civil rights and the Great Society.
Maybe I’m stretching the analogy here, but what claims do you think it is about the American political system that can give you faith that LBJ will produce positive outcomes that you want? What does the LBJ predeployment safety case look like?
Geoffrey Irving: Yeah, so the first thing to say is I’m not going to take a stand at whether he was net good, because he also did a whole bunch of horrible things. I think the take is less that I’m confident that the system in fact aligned him to do good. I think he did probably want to do some good. He just thought, “I must gather all this power along the way to do good,” as many people think.
The case is more that this is a very poorly designed game. A nice analogy, which is fun, which I will cite from that post, is he became Senate majority leader because he realised that position had all this power that everyone else was leaving on the table. For example, he could choose, as majority leader in the Senate, when to call the vote. So he would just sit in the chamber watching people randomly go in and out of the chamber, I don’t know, to the bathroom or to get a snack or something. At some point, the balance of votes in the chamber was in his favour by a few votes — and he would call the vote and win, because he had a perfect memory of who was going to vote for him and extremely good predictions there.
But that is just a very badly designed game that was played. There was this one LBJ guy who’s incredibly good at the details and there was not the competing LBJ force trying to be a counterbalance. So I think when the American system works well, it is because there are effective balances and counterbalances. It’s not clear that those are always working well. But it’s also not clear that the American system is the uniquely best balance/counterbalance system we could have.
We do have the potential to have a more well-designed game and training process, more custom for this process. If you get this kind of counterbalancing, then I think you potentially can get through a lot of the problem.
An example is like if you had the other LBJ that was opposed to the first one, that’s saying, “By the way everyone, you realise what he’s doing here? He’s cheating the vote system.” And everyone is like, “That’s ridiculous. That’s clearly unfair. Let’s fix the rule to break that.” I think that intervention would get you so much power over the misaligned components of LBJ that I think it’s within hope to imagine getting that story right.
Tom Reed: So that’s an example of a system that’s poorly designed but actually reasonably easy to solve.
Tom Reed: What do you think is the role of governments in this world? Resolution’s doing its work. At what point might they need to step in? What might they need to do?
Geoffrey Irving: I think there’s a couple of different levels of government action you could imagine.
Any government can do a bunch of unilateral defensive work. You can work on defences for bio or cyber, or even persuasion potentially. That defensive work can be done by any government kind of unilaterally, and it’s good to do.
Then there’s kind of last-minute temporary pauses, where it’s like, “We’re really close to training this really dangerous model. Let’s chill out for at least a few months and shift resources from capabilities to safety. Try to slow down a little bit, try to just dial up all the knobs that we can in the direction of safety on the margin.” That also means you could, for example, use algorithms which are a significant but not a fatal capability cost hit, like something that’s 2–10x slower. Maybe you can run that in this kind of “temporary pause” world.
Then the more extreme thing is you have a broader treaty where you try to do a longer coordinated slowdown or pause across multiple countries.
I think government should be trying to do all of these things, and then we’ll see how far up the scale we can go. …
My take is that if you were to stop all new model training, there’d be this enormous ongoing wave of economic growth due to the current models. I think if you just take that, it’s enormous in terms of positive benefit, in terms of getting valuable use out of models. You have to learn how to work with the current models, but I think we’re in a massive product overhang. We have worked only a little bit on how to cater to the strengths and weaknesses of models. The models of June 2026 are just incredibly good at software engineering in huge numbers of ways, even before the most recent models in the last couple months.
So I would be fairly unconcerned with that world. It is a tradeoff. I think that if you get stronger models, they can do more things better and probably cheaper. So there’s a tradeoff there. But I think I would much prefer having time to nail down more of the safety story for both alignment and other risks than just massively rolling the dice.
Tom Reed: What is it that gives you so much confidence that we have a high product overhang? If it’s not already showing up in growth statistics, what are the metrics where you’re like, “But look at this thing, it is already very useful, it will lead to lots of economic growth”?
Geoffrey Irving: I think there’s so much use of coding systems in particular, and I think that extends already to huge amounts of other kinds of cognitive labour. Like any kind of analytic analysis of business or the things people can already do with models are so impressive that it is extremely unlikely to me that that has seen kind of full adoption across the economy.
Anecdotally, both from myself playing with models and then just reading a lot about what people are doing, there is a massive learning curve to how to best deploy these models into any particular area of activity. I learn better how to use them across time, and so does everyone else. If we were to stop for even like 10 years, we’ll still keep climbing.
Again, I would be totally lying if I said there wasn’t a tradeoff here. Stronger models are in fact better at doing lots of things, but I would prefer that tradeoff.
Tom Reed: What affordances did you find that you had at UK AISI that you didn’t have at OpenAI or DeepMind for changing the world?
Geoffrey Irving: There’s a couple of them. I’ll list three of them and then we can go from there.
One is adjacency to national security, being close to national security, because there’s a bunch of ingredients out of the risk story that come from those sources, and you need collaborations with natsec to have good takes.
The next one is adjacency to policy. If we want to do this kind of coordination across the world where governments play a role, you sort of have to be in a government to be close to policy in that sense. That’s not the only actor; we want a lot of third parties and nonprofits and independent researchers doing this kind of policy development. But you need part of the story just being in a government.
There’s kind of a subpart of that, which is that in many cases, sometimes governments only listen to governments. At AISI we had a bunch of our own research, but often also we would just be able to go to another government and say, here is some of our research and some of someone else’s research — like from METR or Apollo or the like — and that package was much more received and listened to than if it had just been METR and Apollo trying to go directly to a government of various other countries.
I think that proximity to natsec and policy and other governments of the world is the key thing.
Tom Reed: It’s very valuable. And if there’s so many worlds where governments will need to play a role in things playing out well, what do you think about all the AI researchers who are very concerned about safety, but who are currently working at AI labs rather than in the government? Do you think they’re basically wrong to be doing so?
Geoffrey Irving: Yeah, I think on the margin they are in fact wrong, and many of them should leave and join governments. I think the main argument is that it’s just one of diminishing returns. There are a lot of people at labs. If you are a safety researcher at a lab, probably you’re further out on the diminishing-return curve than you would be if you joined a government or a nonprofit. If every one of the people at labs left en masse and joined the government, that probably would be bad. But that’s not the actual calculation.
Tom Reed: The marginal move is very high value.
Geoffrey Irving: It’s pretty clear. I think people look at themselves and think, “I’m an individual researcher, I’m kind of a special snowflake. I have a very particular agenda, I’m the only one pursuing that particular agenda, I should keep doing it if it’s an important agenda.”
I think that is making a calculation which is a bit too focused, and if you sort of blur your self-image a bit, and just think of it as like, “I’m a safety researcher, I probably have broad takes and knowledge about a variety of things. I can advise governments on a broad range of issues. Probably the lab would pick up the slack on what I’m doing to some degree,” it’ll work pretty well. Again, I think on the margin the calculation is pretty simple.
Tom Reed: What do you think UK AISI specifically will be doing from now until sort of the eve of superintelligence? If they play their hand very well, what kinds of things do you think they’ll be doing that will be moving the needle one way or another?
Geoffrey Irving: Misuse risks are important, so the pure dangerous capability evaluations are important — that story being that high research capability and also close to natsec I think is important for getting those well understood.
Then AISI does a bunch of work on mitigations against both misuse, against loss of control. We have kind of a very strong safeguards team — “we” as in “AISI,” before I left. I think AISI already has strengthened the mitigations of the labs by virtue of being an independent voice and source of research, and that will keep going.
And then the big thing is the main reason I joined the AISI initially: policy. Again, governments have a huge role in policy. AISI is the largest source of government AI research capacity around safety that currently exists, so causing that policy advice to be maximally grounded in the tactical reality of things I think just makes it much more likely to go well.
Tom Reed: Do you think AISI is an asset to the UK specifically? Should every country just have an AISI of its own? How many AISIs do we need?
Geoffrey Irving: I don’t have a confident take there. I think they’re probably more on the margin as good. I think there’s some degree of not wanting to reinvent the wheel too much.
When there are other AISIs, a piece of advice I often give is: it’s important to do a mixture of their own research to build up technical capacity, but then probably don’t try to be a full-on evaluator across all the risks in the same way that [UK] AISI is closer to being. Then be in a position where we can work together across multiple governments, and then to policymakers present: “Here’s all the evidence from all the AISIs plus all the nonprofits kind of appropriately integrated together.” And that, I think, to the extent you can get that kind of collaborative story right, is much more efficient. You get much more knowledge faster across all the governments.
Tom Reed: I’m interested in the version of this world where we do successfully align the superintelligences, we’ve deployed them, and we have high confidence — thanks to Resolution and everyone else’s research — that they will behave the way we want them to. What kind of technologies would you expect that they will develop next?
Geoffrey Irving: All of the practical ones. “Practical” means “allowed by the laws of physics.”
I think we solve ageing, we get nanotech — again, for good or ill; nanotech could be offence- or defence-dominant. Right now software has bugs. Software in the future wouldn’t have bugs, broadly; it would just be perfect in most cases.
I think we will have the ability to colonise the universe in various ways, probably via uploads. We probably will be able to upload humans into machines. My take is that people have this, I think, bad view that the machines will be taking off ahead of us, and then even in the good futures we’ll be stuck behind forever, which I think is wrong. You can imagine uploading someone and then modifying them cognitively — while preserving identity in some meaningful way — to be also superintelligent. So there’s that future ahead of us, should we choose it. Hopefully we have the option to also just live normal lives as humans.
Tom Reed: What happens to the humans that decide not to upload?
Geoffrey Irving: I think they are essentially irrelevant to the economy. But I hope that in this world we will figure out how to derive meaning from family and exploration and so on, whatever the level of cognitive ability is.
Tom Reed: Do you personally expect to upload, by the way?
Geoffrey Irving: Yeah, eventually.
Tom Reed: How would you go about making that decision?
Geoffrey Irving: I don’t think I’d be the first one, but I expect that we’ll just have a good understanding of the science involved. We will have done a bunch of experiments, it will just work very well.
The result is that people will feel great. They’ll be smarter because you can modify them in place in various ways. We’ll understand the brain and AI and so on much better, so that understanding of how to do that modification in a way that is faithful is doable. Yeah, that seems like a good deal.