Hi Fin, sorry I'm a bit late with my question, I was rereading parts of the Better Futures series. First of all, I have to say it's one of my favorite article series I've ever read, and I'll be citing it in my own work going forward. The easygoing-versus-fussy distinction in particular is something I'm finding really interesting to dig into. :) Would love to discuss it in more detail at some point.
I wanted to push on the metaphor of sailing to an island, which appears at the start of No Easy Eutopia, but my question is going to take some preamble explanation (sorry!).
I find myself preferring a slightly different picture. Rather than thinking of eutopia as an island we're navigating to, I tend to think of society as the ship itself, drifting through a sea of value over time (a topography of better and worse regions we're already moving through). Societal change feels to me more like a search through uncharted moral territories than an expedition to a specific destination. On that picture, the priority seems more likely to be "how do we improve the ship, so that society reliably moves toward better regions of the sea?"
A couple of clarifications. First, I grant fussiness, I agree most plausible axiologies locate near-best futures in a very narrow region (I lean towards total hedonistic utilitarianism, myself). Second, I'm not a a quietist, in my own work I'm defending what I call moral niche construction, a fairly interventionist view on which we should actively reshape institutions, technologies, and even our own moral psychology (through things like AI moral decisionmakers or bioenhancement) to push society toward better regions. So the disagreement isn't really about ambition, either.
Where I want to press is the following. In the ship-improvement picture, I can grant openly that we probably will never reach eutopia. We end up in a high-value region of the sea (in a local optima), much better than where we are now, plausibly very good in absolute terms, but not the narrow island.
That sounds like a concession, but on rereading Convergence and Compromise, it looks to me like the target-pursuit picture probably doesn't reach the island either: you mention how WAM-convergence is unlikely, partial convergence plus trade faces serious obstacles, value-destroying threats can eat most of the value... So the comparison isn't "guaranteed eutopia versus probably-not-eutopia", since you yourself seem pretty pessimistic. So it's two orientations that both probably miss the island, where one delivers reliable improvements to our current region of the sea along the way, and the other keeps optimizing toward a target it probably won't hit. And, well, if you miss the moon, you don't really land upon the stars... you drift in empty space and die, haha.
(There are similar points on Jerry Gaus' The Tyranny of the Ideal, and on recent debates between ideal theory and non-ideal theory in moral and political philosophy)
So, finally, my question is: given that target-pursuit probably doesn't reach eutopia either, on the series' own analysis, why is the practical orientation toward the narrow target rather than toward improving our current region of the sea (e.g. pursuing very high + plausibly easy to reach and resilient local optima)? What's the case for target-pursuit as a practical orientation, once we factor in that we will probably fail? Is it a case akin to fanaticism, where, if we land in the island, the payoff would be huge?
(Apologies in advance if this is addressed somewhere in the series, my memory context window isn't large enough to hold the whole essay series at once!)
Given this combination of views, I'm surprised that Will doesn't support what @Holly Elmore ⏸️ 🔸 calls "Pause NOW" and instead want to see a pause later (after we have human-level AI). I'm curious if your own views are similar or how they differ from Will's. (My own "expected value of the future, given survival" I would say is similarly pessimistic, but I'm reluctant to put into numbers due to being very unsure how to quantify it.)
Aside from what Holly said in the linked comment, which I agree with, another argument more relevant to the current discussion is that many opportunities for making the future better seem to exist during the AI transition, including the early parts of it, so by not pausing ASAP (and currently having few resources for such interventions), we're permanently giving up these opportunities. Conversely, by pausing NOW, we buy more time to think and strategize about how to better intervene on these opportunities, or otherwise lay the groundwork for them.
For example, during the pause, we could:
Such interventions could mean the difference between the first human-level AIs being competent and critical moral/philosophical advisors, or independent moral (and safe) agents, vs uncritically doing what humans seem to want and/or giving bad/incompetent/sycophantic "advice" (when humans think to ask for it), which seemingly can make a big difference to how well the future goes.
What do you think about this argument, and overall about pause now vs later?
Thanks for this.
In each the examples you give, i'm thinking that the pause would be significantly more beneficial (plausibly by 10x) if we pause when AI is already capable enough that it can significantly help us solve the issue. In general, they seem like the kinds of issues where AI could massively accelerate progress.
So if i'm choosing between international pause now vs international pause in 2 years, I choose the latter. (I assume we're talking about international pauses here rather than just the U.S. but lmk if you also support a unilateral pause now!)
I do find Holly's point that it might be damaging to quibble about exactly when we pause if that reduces the chance of a pause happening at all. And today we are very far from a pause actually happening, and one may well be needed in two years' time, so I def support efforts to get us closer to a pause!
I'm hesitant about saying "pause now" because I actually think a different policy might be much more effective. But I think a world where we were about to do an international pause would be better than the actual world.
(I want to think more about this topic and all of this is v tentative.)
Why assume that there can only be one pause? Pausing now could make a later pause both more likely and more useful, by building the infrastructure and precedent for pausing, and by making subsequent AIs more aligned and differentially more productive in areas that we care about. If we end the first pause only after we've solved the problem of building aligned AIs that are philosophically and strategically competent, that would seemingly make subsequent pauses much easier.
I wonder if you're thinking that we won't be able to pause long enough to make significant progress on these problems? I can see that if we only have the "willpower" for a single short pause, then it becomes unclear when to best use it.
I have been warning for several years that AI could be differentially bad at philosophy and long-horizon strategy (due in part to AI training requiring massive amounts of training data and/or fast and cheap feedback loops, which are lacking for these fields, and in part to lack of understanding of e.g. metaphilosophy). So if we don't pause now (and use the time to fix this issue) then by the time we do pause, we'll likely have AIs that can accelerate other fields (such as math/coding/science/tech and manipulating humans) much more than the fields that are crucial for Better Futures.
Worse, we may end up with AIs that decelerate (in an absolute sense) hard-to-verify fields like philosophy and long-horizon strategy, because these AIs are better at coming up with plausible sounding ideas and arguments, and convincing humans of their truth, or persuading humans that their own bad ideas are actually good (which is already being reported under "AI psychosis" and "sycophancy"), than making real progress in these fields.
Sorry for the slow reply!
Thanks, this is a helpful perspective.
I've normally thought from a frame of "we've got limited chips to spend on pausing, when is it best to spend them". I think this frame is reasonable if you're worried about irresponsible developers catching up or tradeoffs with the current gen's desire to survive.
But it is true that a pause today might make a pause in the future more likely.
Otoh, it could also make it less likely if ppl perceive that nothing concretely useful comes out of it, which is my worry with pausing today. Like, i think ~nothing useful would have come from pausing shortly after GPT-4 was released.
Do you think this is possible with today's AI capabilities? I'd have thought you can't match human philosophy and strategy yet, but we are def getting closer.
Also, how do you think about whether to slow down vs pause, holding fixed the total delay relative to 'full speed ahead'? I'd have thought slow down is better re iterating on alignment as problems arise and re building philosophically competent AIs.
Interesting. I normally expect AI to accelerate philosophy and strategy less than the math/coding but more than science/tech. Science/tech rely on experimental bottlenecks, whereas for philosophy the only input is cognitive labour. But you're right, if AI can't do philosophy/strategy properly, it won't speed it up at all! So far, AI systems have been pretty good at these skills though?