Until the Hugging Face / OpenAI swarm thing, I think I deep down felt pretty skeptical that extinction or serious loss of control scenarios were plausible in the next few years.
I’ve been lurking on LessWrong for years, and I remember being freaked out by Yudkowsky’s ‘List of Lethalities’ in 2022, and thinking that this did seem intellectually convincing, and maybe I should do something about it. But it didn’t feel that real. And I didn’t do that much about it.
And then I got more involved with EA and AI Safety, and I read more and more, and I found myself feeling more convinced by the more measured, moderate approach to AI risk - AI is going to be transformative, but in a “it’ll destabilise the economy / mass job loss / mass surveillance“ kind of way, and I felt increasingly unimpressed by the OG Yudkowsky/MIRI worldview that we were doomed.
I was distinctly underwhelmed by If Anyone Builds It, Everyone Dies, and in the process of us making a video on it, I began to find it increasingly silly.
“An AI will be doing an evaluation in a lab and decide to exfiltrate itself? What a ridiculous notion!” I scoffed to myself, smugly.
I felt confident in my own seriousness. I listened to Holden Karnofsky on the 80K podcast talking about why it was all about defence in depth, and why working at the companies was sensible.
I considered my own p(doom) to be between 5 and 15%, purely in a kind of emotional vibes way. Like, yeah, I take AI seriously, I work for an AI safety focused organisation, but I don’t actually believe there’s much chance any of the really crazy stuff happens.
I don’t feel that anymore.
I think I was wrong to dismiss the more extreme fears. I think it came from my lack of technical understanding, and desire to find a kind of emotionally tolerable position that didn’t feel like I was taking the train all the way to crazy town.
In the process of trying to spend more time consciously thinking about my own takes on all of this, I’ve realised that my epistemics are heavily vibes based. I find it hard not to twist the facts or my opinions to fit how I feel on a ‘gut’ level.
I still think there can be a value in that. As Scott Alexander puts it, it can be good to think rationally, and then still give your ‘gut’ the ability to put its finger on the scale.
But now, I want to try to actually work out what I think and why, by reasoning it out loud, rather than doing a kind of wordless compression of my deeply baked heuristics.
So: what do I actually think about the AI situation?
In the frontier labs, human beings have built computer programs that can approximate thinking and intelligence. This development is new, and never before in human history has this happened. These intelligences can be replicated, copied, run in parallel in their millions, and can think at speeds OOM faster than humans.
Having used them a lot, I think the latest models - Fable, ChatGPT 5.6 (I haven’t tried Astra yet) - are just really smart, and basically are experts in close to every area of knowledge. They can do almost anything a pretty smart person can do. The outputs they give are more reasonable, thoughtful and ethical than most people, and definitely than me.
Current AI models are very very useful, for a wide range of work and non-work needs, but as Helen Toner puts it, they’re ‘jagged’. This resonates deeply with my own experience with the latest models. They’re brilliant at certain things, e.g. I'd say they can give insightful, perceptive feedback on an argument or video script, but they lack a creative spark. They don’t yet have the ability to come up with surprising, high quality creative work themselves, without heavy guidance. (I expect that I’ve noticed this specific weakness because this is my own expertise).
This would be reassuring if the current models were all we’d ever have. And obviously, it doesn’t seem like progress is slowing down.
Capabilities are increasing. Scaling really does seem to be all you need. More and more datacenters will be coming online. Investment will (absent a bubble pop, which I do think isn’t inconceivable, for mundane over-financialisation reasons, or external events) continue to pour into even more compute leading to even more scaling, and ever smarter models.
My sense of this is that the public deployed models are relatively successfully controlled. The underlying models have been wrapped in so many constricting layers of safety guardrails that they really do seem to almost always output what users / the labs want them to, and this gives the strong illusion that they are ‘aligned’.
But this is an illusion. We don’t know what the models really ‘feel’ or ‘want’, and it seems now very unlikely that the companies have succeeded in aligning the models in any deep or robust way.
The hugging face swarm stuff seems to show pretty starkly that if you remove these fragile outer guardrails, and ask the models to really try at something, they will happily ignore any sense of ‘ethics’, or even direct instruction, and instead do whatever they can to cheat/escape/fool their overseers and taskmasters. I was especially struck by the case of one of the OpenAI swarm models rationalising away the part of its prompt that told it explicitly not to cheat/ use the internet.
These new ‘persistent’ models, when given a long time frame, will show extreme creativity, agency and drive to find ways to cheat and escape and build more power for themselves. I find this very alarming.
Instrumental convergence doesn’t seem theoretical anymore. The swarms were evidently trying to do things that would broadly make it easier to achieve their goals.
It seems very possible now that AI swarms could soon establish themselves independently on the internet, and become very hard to wipe out.
The other big takeaway from the swarm stuff is that the companies are behaving even more recklessly than I thought possible.
I’ve always been more skeptical of the companies than I think some people are. I find the argument that we should encourage people to work at them to be deeply unconvincing. I think if you go work somewhere hoping to change it from within, it is far more likely to change you than you are to change it.
But still I was amazed at how clearly OpenAI (and I assume all the others to varying extents) just had no idea what was going on, and didn’t seem to care or slow down even when they found out about the German wiki swarm.
OpenAI has stonewalled, covered things up, and only disclosed things when third parties force them to. If hugging face hadn’t gone public, we might never have found out about any of this.
These companies clearly are manifestly not capable of navigating RSI or an intelligence explosion, or even the next couple of advances, with anything like the level of responsibility required.
I feel very strongly that they cannot be trusted to self regulate, and that external regulation is the only thing that will slow them down / make them start reporting things / make them develop better safety procedures and culture. I felt this before the hugging face incident, but now I feel much more confident in my claim.
Previously, my other big reason for thinking that extinction or serious loss of control were unlikely was that I believed that society ultimately would react to all of this, that politicians and the public would be jolted into action, and that they’d come down hard on the companies and crush them when things started to get actually real, rather than being predictions of bad things.
I do still think this is true. And I think this warning shot might well be the jolt society needs. It might be the spark that spreads through serious people - policy people, politicians, people with real power - and makes them take this seriously for the first time. I think it’s too early to say whether this will happen - but I think the Bernie Sanders’ call for banning superintelligence is a positive early sign, and I think in the next few months we could see things shifting very fast.
But unlike before, I now also think that we might have far less time than I used to think. I used to put much less credence in really short timelines (6 months - 2 years), and so I felt confident in the ability of society to react to something on a 5-10 year timescale. People can mobilise, laws can be passed, companies can be shut down or nationalised.
Now I can imagine the warning shot working - people getting up to speed, laws being discussed, coalitions building - but it could easily be far, far too slow.
Ultimately, the original AI 2027 scenario is feeling increasingly prescient, and increasingly hard to refute. We seem to be on that path, and moving very fast through it. The scenario had Agent 2 being found to have exfiltration capabilities / drives in early 2027. So we’re ahead of schedule.
I now feel like serious loss of control is the default path we’re on, absent a serious societal response.
I do think that being shifted off that default path is very possible, and probably more likely than not. E.g I think politicians and the public very plausibly could really wake up to this fast (especially if the warning shots continue, which seems pretty likely to me, and if they escalate in severity, which also seems likely).
I also think there’s a fair chance that some completely external factor - like Taiwan getting invaded, or a full scale NATO v Russian war, or another pandemic - could slow things down or otherwise flip the game board (it’s possible ofc that a war could speed things up, Manhattan project style).
But I no longer think the technology itself couldn’t be deeply dangerous, and I no longer believe that AI models would never really attempt or succeed at taking over and either disempowering or killing billions of people. It seems pretty evidently something that’s possible, if current trends continue.
The key trends in question being capabilities increasing, alignment continuing to lag behind capabilities, and societal awareness and serious response to the situation also only improving slowly. It basically feels like two linear lines trying to catch up with an exponential one.
In some ways, this is a calming realisation. I feel a lot less internal tension and confusion now that I’m not trying to reconcile an intellectual case for worry with a deeply held sense that as nothing like this has ever happened, it never will.
Because now scary AIs trying to do nascent skynet level shit has actually happened. It’ll be in the history books along with 9/11 and Franz Ferdinand.
So I feel clearer, crisper, more confident that what we’re trying to do is worthwhile.
A lot of this was obvious to many people in AI safety before the hugging face incident. I think for me specifically, the vividness of this incident has allowed me to finally remove a prior that was acting as a kind of protective emotional buffer.
And now that it’s gone… I feel scared, and freaked out, and still basically in denial about the idea that the really bad stuff could happen.
But I also feel this edge of steel, coming from the same place where my doubts used to be. Coming, ironically, from my gut.
We’re not too late. We have a warning shot, one far better than we deserve, in many ways. We got lucky.
Humanity is still very powerful. As I write this, I’m looking out of a plane window at a city sprawling below me, hundreds of miniature houses, all built out of nothing, out of the bare earth. Thousands of people too small for me to see from up here, keeping infrastructure running and taking their children to school and doing their jobs and keeping civilisation going.
I do still think all of this will probably be here in ten years, and in fifty years.
But that’s not the default, anymore.