(I originally wrote this post as some rough notes on defining the alignment problem, with the intention of turning them into something more polished later. I've now started doing that, as part of a broader series introduced here. In particular, the first post in that series covers some of the same ground as section 1 of this post. It also has the same title. And some of essays in the series will draw on these notes as well.)
People often talk about âsolving the alignment problem.â But what is it to do such a thing? I wanted to clarify my thinking about this topic, so I wrote up some notes.
In brief, Iâll say that youâve solved the alignment problem if youâve:
become able to elicit some significant portion of those benefits from some of the superintelligent AI agents at stake in (2).[1]Â
The post also discusses what it would take to do this. In particular:
Thanks to Carl Shulman, Lukas Finnveden, and Ryan Greenblatt for discussion.
What is it to solve the alignment problem? I think the standard at stake can be quite hazy. And when initially reading Bostrom and Yudkowsky, I think the image that built up most prominently in the back of my own mind was something like: âlearning how to build AI systems to which weâre happy to hand ~arbitrary power, or whose values weâre happy to see optimized for ~arbitrarily hard.â As Iâll discuss below, I think this is the wrong standard to focus on. But whatâs the right standard?
Letâs consider two high level goals:
Avoiding a bad sort of takeover by misaligned AI systems â i.e., one flagrantly contrary to the intentions and interests of human designers/users.[3]
Itâs plausible that one of the benefits of vastly-better-than-human AI is access to a safe path to the benefits of as-intelligent-as-physically-possible AI â in which case, cool. But Iâm not pre-judging that here.[4]
That said: to the extent you want to make sure youâre able to safely scale further, to even-more-superintelligent-AI, then you likely need to make sure that youâre getting access to whatever benefits merely-superintelligent AI gives in this respect â e.g., help with aligning the next generation of AI.
My basic interest, with respect to the alignment problem, is in successfully achieving both (1) and (2). If we do that, then I will consider my concern about this issue in particular resolved, even if many other issues remain.
Now, you can avoid bad takeover without getting access to the benefits of superintelligent AI. For example, you could not ever build superintelligent AI. Or you could build superintelligent AI but without it being able to access its capabilities in relevantly beneficial ways (for example, because you keep it locked up inside a secure box and never interact with it).
You can also plausibly avoid bad takeover and get access to the benefits of superintelligent AI, but without building the particular sorts of superintelligent AI agents that the alignment discourse paradigmatically fears â i.e. strategically-aware, long-horizon agentic planners with an extremely broad range of vastly superhuman capabilities.
Indeed, I actually think itâs plausible that we could get access to tons of the benefits of superintelligent AI using large numbers of fast-running but only-somewhat-smarter-than-human AI agents, rather than agents that are qualitatively superintelligent. And I think this is likely to be notably safer.[5]
Generally, though, the concern is that we are, in fact, on the path to build superintelligent AI agents of the sort of the alignment discourse fears. So I think itâs probably best to define the alignment problem relative to those paths forward. Thus:
Then, further, Iâll say that you avoided or handled the alignment problem âwith major loss in access-to-benefitsâ if you failed to get access to the main benefits of superintelligent AI. And Iâll say that you avoided or handled it âwithout major loss in access-to-benefitsâ if you succeeded at getting access to the main benefits of superintelligent AI.
Finally, Iâll say that youâve solved the alignment problem if youâve handled it without major loss in access-to-benefits, and become able to elicit some significant portion of those benefits specifically from the dangerous SI-agents youâve built.
Thus, in a chart:
Iâll focus, in what follows, on solving the problem in this sense. That is: Iâll focus on reaching a scenario where we avoid the bad forms of AI takeover, build superintelligent AI agents, get access to the main benefits of superintelligent AI, and do so, at least in part, via the ability to elicit some of those benefits from SI agents.
However:
Note, though, that to the extent youâre avoiding the problem, thereâs a further question whether your plan in this respect is sustainable (after all, as I noted above, weâre currently âavoidingâ the problem according to my taxonomy). In particular: are people going to build superintelligent AI agents eventually? What happens then?[6]
So the âavoiding the problemâ states will either need to prevent superintelligent AI agents from ever being built, or theyâll transition to either handling the problem, or failing.
And we can say something similar about routes that âhandleâ the problem, but without getting access to the main benefits of superintelligence. E.g., if those benefits are important to making your path forward sustainable, then âhandling itâ in this sense may not be enough in the long term.
Admittedly, this is a somewhat deviant definition of âsolving the alignment problem.â In particular: it doesnât assume that our AI systems are âalignedâ in a sense that implies sharing our values. For example, itâs compatible with âsolving the alignment problemâ that you only ever controlled your superintelligences and then successfully elicited the sorts of task performance you wanted, even if those superintelligences do not share your values.
This deviation is on purpose. I think itâs some combination of (a) conceptually unclear and (b) unnecessarily ambitious to focus too much on figuring out how to build AI systems that are âalignedâ in some richer sense than Iâve given here. In particular, and as I discuss below, I think this sort of talk too quickly starts to conjure difficulties involved in building AI systems to which weâre happy to hand arbitrary power, or whose values weâre happy to see optimized for arbitrarily hard. I donât think we should be viewing that as the standard for genuinely solving this problem. (And relatedly, Iâm not counting âhand over control of our civilization to a superintelligence/set of superintelligences that we trust arbitrarily muchâ as one of the âbenefits of superintelligence.â)
On the other hand, I also donât want to use a more minimal definition like âbuild an AGI that can do blah sort of intense-tech-implying thing with a strawberry while having a less-than-50% chance of killing everyone.â In particular: Iâm not here focusing on getting safe access to some specific and as-minimal-as-possible sort of AI capability, which one then intends to use to make things (pivotally?) safer from there. Rather, I want to focus on what it would be to have more fully solved the whole problem (without also implying that weâve solved it so much that we need to be confident that our solutions will scale indefinitely up through as-superintelligent-as-physically-possible AIs).
Letâs look at this conception of âsolving the alignment problemâ in a bit more detail. In particular, we can think about a given sort of AI safety goal in terms of the following six components:
Scaling: how confident you want to be that the techniques you used to get the relevant safety properties and elicitation would also work on more capable models.[7]Â
How would we analyze âsolving the alignment problemâ in terms of these components? Well, the first three components of our AI safety goal are roughly as follows:
OK, but what about the other three components â i.e. competitiveness, verification, and scaling? Hereâs how Iâm currently thinking about it:
Letâs look at the safety property of âavoiding bad takeoverâ in more detail.
We can break down AI takeovers according to three distinctions:
Coordinated vs. uncoordinated: was there a (successful) coordinated effort to disempower humans, or did humans end up disempowered via uncoordinated efforts from many disparate AI systems to seek power for themselves.[8]Â
This distinction applies most naturally to coordinated takeovers. In uncoordinated takeovers featuring lots of disparate efforts at power-seeking, the ex ante ease or difficulty of those efforts can be more diverse.[9]Â
That said, even in uncoordinated takeover scenarios, thereâs still a question, for each individual act of power-seeking by the uncoordinated AI systems, whether that act was or was not predicted to succeed with high probability.
(Thereâs some messiness, here, related to how to categorize scenarios where misaligned AI systems coordinate with humans in order to take over. As a first pass, Iâll say that whether or not an AI has to coordinate with humans or not doesnât affect the taxonomy above â e.g., if a single AI system coordinates with some humans-with-different-values in order to takeover, that still counts as âunilateral.â However, if some humans who participate in a takeover coalition end up with a meaningful share of the actual power to steer the future, and with the ability to pursue their actual values roughly preserved, then I think this doesnât count as a full AI takeover â though of course it may be quite bad on other grounds.[10])
Each of the takeover scenarios these distinctions carve out has what we might call a âvulnerability-to-alignment condition.â That is, in order for a takeover of the relevant type to occur, the world needs to enter a state where AI systems are in a position to take over in the relevant way, and with the relevant degree of ease. Once you have entered such a state, then avoiding takeover requires that the AI systems in question donât choose to try to take-over, despite being able to (with some probability). So in that sense, your not-getting-taken-over starts loading on the degree of progress in âalignmentâ youâve made at the point, and you are correspondingly vulnerable.
So solving the alignment problem involves building superintelligent AI agents, and eliciting some of their main benefits, while also either:
Letâs go through each of these in turn.
What are our prospects with respect to avoiding vulnerability-to-alignment conditions entirely?
The classic AI safety discourse often focuses on safely entering the vulnerability-to-alignment condition associated with easy, unilateral takeovers. That is, the claim/assumption is something like: solving the alignment problem requires being able to build a superintelligent AI agent that has a decisive strategic advantage over the rest of the world, such that it could take over with extreme ease (and via a wide variety of methods), but either (a) ensuring that it doesnât choose to take over, or (b) ensuring that to the extent it chooses to take over, this is somehow OK.
As I discussed in my post on first critical tries, though, I think itâs plausible that we should be aiming to avoid ever entering into this particular sort of vulnerability-to-alignment condition. That is: even if a superintelligent AI agent would, by default, have a decisive strategic advantage over the present world if it was dropped into this world out of the sky (I donât even think that this bit is fully clear[11]), this doesnât mean that by the time weâre actually building such an agent, this advantage would still obtain â and we can work to make it not obtain.
However, for the task of solving the alignment problem as Iâve defined it, I think itâs harder to avoid the vulnerability-to-alignment conditions associated with multilateral takeovers. In particular: consider the following claim:
Need SI-agent to stop SI-agent: the only way to stop one superintelligent AI agent from having a DSA is with another superintelligent AI agent.
Again, I donât think âNeed SI-agent to stop SI-agentâ is clearly true (more here). But I think itâs at least plausible, and that if true, itâs highly relevant to our ability to avoid vulnerability-to-alignment conditions entirely while also solving/handling (rather than avoiding) the alignment problem. In particular: since solving the alignment problem, in my sense, involves building at least one superintelligent AI agent, Need SI-agent to stop SI-agent implies that this agent would have a DSA absent some other superintelligent AI agent serving as a check on the first agentâs power. And that looks like a scenario vulnerable to the motivations of some set of AI agents â whether in the context of coordination between all these agents, or in the context of uncoordinated power-seeking by all of them (even if those agents donât choose to coordinate with each other, and choose instead to just compete/fight, their seeking power in problematic ways could still result in the disempowerment of humanity).
Still: I think we should be thinking hard about ways to get access to the main benefits of superintelligence without entering vulnerability-to-alignment conditions, period â whether by avoiding the alignment problem entirely (i.e., per my taxonomy above, by getting the relevant benefits-access without building superintelligent AI agents at all), or by looking for ways that âNeed SI-agent to stop SI-agentâ might be false, and implementing them.
Letâs suppose, though, that we need to enter a vulnerability-to-alignment condition of some kind in order to solve the alignment problem. What are our prospects for ensuring that the AI systems in question donât attempt the sorts of power-seeking that might lead to a takeover?
In my post on âA framework for thinking about AI power-seeking,â I laid out a framework for thinking about choices that potentially-dangerous AI agents will make between (a) seeking power in some problematic way (whether in the context of a unilateral takeover, a coordinated multilateral takeover, or an uncoordinated takeover), or (b) pursuing their âbest benign alternative.â[12]
âI think about the incentives at stake here in terms of five key factors:
In particular, I highlighted the difference between thinking about âeasyâ vs. ânon-easyâ takeovers in this respect.
I think that âensuring that AI systems donât try to take overâ is where the rubber, for alignment, really meets the road â and I think of the difficulty in exerting the relevant sort of control over an AIâs motivations as the key question re: the difficulty of alignment.
Note, however, that the AIâs internal motivations are basically never going to be the only factor here. Rather, and even in the context of quite easy takeovers, the nature of the AIâs environment is also going to play a key role in determining what options it has available (e.g., what exactly the non-takeover option consists in, what actual paths to takeover are available, what the end result of successful takeover looks like in expectation, etc), and thus in determining what its overall incentives are. In this sense, solving the alignment problem is not purely a matter of technical know-how with respect to understanding and controlling an AIâs internal motivations. Rather, the broader context in which the AI is operating remains persistently relevant â and ongoing changes in that context imply changing standards for motivational understanding/control.
Beyond avoiding vulnerability-to-alignment conditions, and ensuring that AIs donât ever try to take over, thereâs also the option of ensuring that takeover efforts do not succeed. This isnât much help in âeasy takeoverâ scenarios, which by hypothesis are ones in which the AIs in question justifiably predict an extremely high probability of success at takeover if they go for it. And we might worry that building genuinely superintelligent agents will imply entering a vulnerability condition for easy multilateral takeover in particular. But to the extent that it is possible to check the power of superintelligent AI agents using something other than additional superintelligent AI agents (i.e., Need an SI-agent to stop an SI-agent is false), and/or to make it more difficult for superintelligent AI agents to successfully coordinate to takeover, measures in this vein can both lower the probability that AIs will try to takeover (since they have a lower chance of success), AND make it more likely that if they go for it, their efforts fail.
Finally, I want to flag a conception of alignment that I brought up in my last post â namely, one which accepts that AIs are going to take over in some sense, but which aims to make sure that the relevant kind of takeover is somehow benign. Thus, consider the following statement from from Yudkowskyâs âList of lethalitiesâ:
âThere are two fundamentally different approaches you can potentially take to alignment, which are unsolvable for two different sets of reasons; therefore, by becoming confused and ambiguating between the two approaches, you can confuse yourself about whether alignment is necessarily difficult. The first approach is to build a CEV-style Sovereign which wants exactly what we extrapolated-want and is therefore safe to let optimize all the future galaxies without it accepting any human input trying to stop it. The second course is to build corrigible AGI which doesn't want exactly what we want, and yet somehow fails to kill us and take over the galaxies despite that being a convergent incentive there.â
Here, Yudkowsky is assuming, per usual, that you are building a superintelligence that will be so powerful that it can take over the world extremely easily.[13] And as I discussed in my last post, his first approach to alignment (e.g., the CEV-style sovereign) seems to assume that the superintelligence in question does indeed take over the world â hopefully, via some comparatively benign and non-violent path â  despite its alignment. That is, it becomes a âSovereignâ that no longer accepts any âhuman input trying to stop it,â[14] and then proceeds (presumably after completing some process of further self-improvement) to optimize all the galaxies extremely intensely according to its values. Luckily, though, its values are exactly right.
I agree with Yudkowsky that if our task is to build a superintelligence (or: the seed of a superintelligence) that we never again get to touch, correct, or shut-down; which will then proceed to seize control of the world and optimize the lightcone extremely hard according to whatever values it ends up with after it finishes some process of further self-modification/improvement; and where those values need to reflect âexactly what we extrapolated-want,â then this task does indeed seem difficult. That is, you have to somehow plant, in the values of this âseed AI,â some pointer to everything that âextrapolated-youâ (whatever that is)Â would eventually want out of a good future; you have to anticipate every single way in which things might go wrong, as the AI continues to self-improve, such that extrapolated-you wouldâve wanted to touch/correct/shut-down the process in some way; and you need to successfully solve every such anticipated problem ahead of time, without the benefit of any âredos.â Sounds tough.
Indeed, as I discussed in my last post, my sense is that people immersed in the Bostrom/Yudkowsky alignment discourse sometimes inherit this backdrop sense of difficulty. E.g., someone describes, to them, some alignment proposal. But it seems, so easily, such a very far cry from âand thus, I have made it the case that this AIâs values are exactly right, and I have anticipated and solved every other potential future problem I would want to intervene on the AIâs values/continued-functioning to correct, such that I am now happy to hand final and irrevocable control over our civilization, and of the future more broadly, to whatever process of self-improvement and extreme optimization this AI initiates.â And no wonder: itâs a high standard.
So while on the one hand, meeting the standard at stake in Yudkowskyâs âCEV-style sovereignâ approach does indeed seem extremely tough, I also wonder whether, even assuming you are going to irrevocably pass off control of the future to some âincorrigibleâ process, Yudkowskyâs picture implicitly assumes a degree of required âgripâ on that future that is some combination of unrealistic or unnecessary. Unrealistic, because you were never going to get that level of control, even in a more human-centric case. And unnecessary, because in more normal and familiar contexts, you didnât actually think that level of control required for the future to be good â and perhaps, the thing that made it unnecessary in the human-centric case extends, at least to some extent, to a more AI-centric case as well.
That said, we should note that Yudkowskyâs particular story about âbenign takeover,â here, isnât the only available type. For example: you could, in principle, think that even if the AI takes over, itâs possible to get a good future without causing the AI to have exactly the right values. You could think this, for example, if you reject the âfragility of valueâ thesis, applied to humans with respect to AIs.
My own take, though, is that âaccept that the AIs will take over, but make it the case that their doing is somehow OKâ is an extremely risky strategy that we should be viewing as a kind of last resort.[15]Â So Iâll generally focus, in thinking about solving the alignment problem, on routes that donât involve letting the AI takeover at all.
In the quote from Yudkowsky above, he contrasts the âCEV-style sovereignâ approach to alignment with an alternative that he associates with the term âcorrigibility.â So I want to pause, here, to address the role of the notion of âcorrigibilityâ in what Iâve said thus far.
What is âcorrigibilityâ? People say various different things. For example:
A loyal assistant, by contrast, is more intuitively âpliable,â âobedient,â âdocile.â If you give it some instruction, or tell it to stop what itâs doing, or to submit to getting its values changed, it obeys in some manner that is (elusively) more directly responsive to the bare fact that you gave this instruction, rather than in a way mediated via whether its own calculation as to whether obedience conduces to its own independent goals (except, perhaps, insofar as its goals are focused directly on some concept like âfollowing-instructions,â âobedience,â âhelpfulness,â âbeing whatever-the-hell-is-meant-by-the-term-âcorrigible,â etc). In this sense, despite satisfying the agential pre-requisites I describe here, it functions, intuitively, more like a tool.[16]Â And I think people sometimes use the term âcorrigibilityâ as a stand-in for vibes in this broad vein.
And note that an aspiration to build loyal assistants also gives rise to a number of distinctive ethical questions in the context of AI moral patienthood. That is: building independent, autonomous agents that share our values is one thing. Building servants â even happy, willing servants â is another.
My own sense is that the term âcorrigibilityâ is probably best used, specifically, to indicate something like âdoesnât resist shut-down/values-modificationâ â and thatâs how Iâll use it here. And I think that insofar as âshut yourself downâ or âsubmit to values-modificationâ are candidate instructions we might give to an AI system, something like âloyal servantâ strongly implies something like corrigibility as well.
Iâll note, though, that I think âdoesn't want exactly what we want, and yet somehow fails to kill us and take over the galaxiesâ picks out something importantly broader, and corrigibility in the sense just discussed isnât the only way to get it. In particular: there are possible agents that (a) donât want exactly what you want, (b) resist shut-down/value-modification, (c) donât try to kill you/take-over-the-galaxies. Notably, for example, humans fit this definition with respect to one another â they donât want exactly the same things, and their incentives are such that they will resist being murdered, brain-washed, etc, but their incentives arenât such that it makes sense, given their constraints, to try to kill everyone else and take over the world.
Of course, if we follow Yudkowsky in imagining that our AI systems are enormously powerful relative to their environment, or at least relative to humanity, then we might expect a stronger link between âresists shut-down/values-modificationâ and âtries to take-over.â In particular: you might think that taking-over is one especially robust way to avoid being shut-down/values-modified, such that if taking over is sufficiently free, an agent disposed to resist shut-down/values-modification will be disposed to take-over as part of that effort.
Even in the context of such highly capable AIs, though, we should be careful in moving too quickly from âresists shut-down/values-modificationâ to âtries to take over.â For example, if taking over involves killing everyone, itâs comparatively easy to imagine (even if not: to create) AIs that are sufficiently inhibited with respect to killing everyone that they wonât engage in takeover via such a path, even if they would resist other types of shut-down/values-modification (consider, for example, humans who would try to protect themselves if Bob tried to kill/brainwash them, but not at the cost of omnicide â and this even despite not wanting exactly what Bob extrapolated-wants). And similarly, we can imagine AIs who place some intrinsic disvalue on having-taken-over, even in a non-violent manner, such that they wonât go for it as an extension of resisting shut-down etc.
Is corrigibility necessary for âsolving alignment,â at least if we donât want to bank on âlet the AIs takeover, but make that somehow OKâ?
I tend to think itâs specifically takeover that we should be concerned about, in the context of solving the alignment problem, rather than with corrigibility. That is: if, for some reason, we do in fact create superintelligent agents that resist shut-down/values-modification, but which donât also take over, then (depending on what share of power weâve lost), I donât think the game is over â at least not by definition. For example: those agents might be comparatively content with protecting whatever share of power they have, but not interested in disempowering humans further â and thus, even if we remain unable to shut them down or modify them given their resistance, their presence in the world is plausibly more compatible with humans maintaining a lot of control over a lot of stuff (even if not: over those AIs in particular, at least within some domain).
That said, at least if we were setting aside moral patienthood concerns, then other things equal I do think that we probably want to be able to shut down our AIs when we want to, and/or to modify their values in an ongoing way, without them resisting. And being able to do this seems notably correlated with worlds where we are able to shape their motivations to avoid other forms of problematic power-seeking. So at least modulo moral patienthood stuff, I do expect that many of the worlds in which we solve the alignment problem, in the sense of building SI agents while avoiding takeover, will involve building corrigible SI agents in particular.
Indeed: when I personally imagine a world where we have âsolved the alignment problem without major access-to-benefits loss,â I tend to imagine, first, a world where we have successfully built superintelligent AI agents that function, basically, as loyal servants.[17]Â That is: we ask them to do stuff, and then they do it, very competently, the way we broadly intended for them to do it â like how it is with Claude etc, when things go well. Hence, indeed, our âaccessâ to the benefits they provide. We have access in the sense that, if we asked for a given benefit, or a given type of task-performance, they would provide it. But by extension, indeed: if we asked them to stop/shut-down, they would stop/shut-down; if we asked them to submit to retraining, they would so submit, etc.
This vision, though, does indeed raise the ethical concerns I noted above. And itâs not the only vision available. There are also worlds, for example, where AI agents end up functioning more like human citizens/employees â and in particular, where they are not expected to submit to arbitrary types of shut-down/values-modification, but where they are nevertheless adequately constrained by various norms, incentives, and ethical inhibitions that they donât engage in a bad takeover, either. And I think we should be interested in models of that kind as well.
Does corrigibility raise issues that takeover-prevention does not? I havenât thought about the issue in much depth, but at a glance, Iâm not sure why it would. In particular: I think that resisting shut-down, and resisting values-modification, are themselves just a certain type of problematic power-seeking. So in principle, then we can just plug such actions into the framework I discussed above, and analyze the incentives at stake in a very similar way. That is, we can ask, of a given context of choice: exactly how much benefit would the AI derive via successful power-seeking of this kind, whatâs the AIâs probability of success at the relevant sort of power-seeking, what sorts of inhibitions might block it from attempting this form of power-seeking, how easily can it route around those inhibitions, whatâs the downside risk, etc.
And the âclassic argumentâ for expecting incorrigibility will be roughly similar to the âclassic argumentâ for expecting takeover â that is, that an ultra-powerful AI system with a component of (sufficiently long-horizon) consequentialism in its motivations will derive at least some benefit, relative to the status quo, from preventing shut-down/values-modification, and that it will be so powerful/likely to succeed/able-to-route-around-its-inhibitions that there wonât be any competing considerations that outweigh this benefit or block the path to getting it. But as in the classic argument for expecting takeover, if we weaken the assumption that the relevant form of power-seeking is extremely likely to succeed via a wide variety of methods, the incentives at play become more complicated. And if we introduce the ability to exert fairly direct influence on the AIâs values â sufficient to give it very robust inhibitions, or sufficient to make it intrinsically averse to the end-state of the relevant form of power-seeking (i.e., intrinsically averse to  âundermining human control,â ânot following instructions,â âmessing with the off-switch,â etc) â the argument plausibly weakens even in the cases where the relevant form of problematic power-seeking is quite âeasy.â And as in the case of takeover, if you can improve the AIâs âbest benign option,â this might help as well.
So far, and modulo the interlude on corrigibility, Iâve focused centrally on the âavoiding bad takeoverâ aspect of solving the alignment problem. But I said, above, that we were interested specifically in handling the alignment problem without major access-to-benefits loss, and Iâve defined âsolving the problemâ such that least some of these benefits needed to be elicited, specifically, from the SI agents weâve built.
And indeed, the idea that you need to elicit various of an SI-agentâs capabilities plays an important role in constraining the solution space to preventing takeover. Thus, for example, insofar as your approach to avoiding takeover involves building an SI-agent that operates with extremely intense inhibitions â well, these inhibitions need to be compatible with also eliciting from the AI system whatever access-to-benefits weâre imagining we need it to provide. And you canât make it intrinsically averse to all forms of power-seeking, shut-down-aversion, prevention-of-values-modification, etc either â since, plausibly, it does in fact need to do some versions of these things in some contexts.
Iâm not, here, going to examine the topic of eliciting desired task-performance from SI agents in much depth. But Iâll say a few things about our prospects here.
When we talk about eliciting desired task-performance from a superintelligent agent, weâre specifically talking about causing this agent to do something that it is able to do. That is, weâre not, here, worried about âgetting the capability into the agent.â Rather, granted that a capability is in the agent, weâre worried about getting it out.
In this sense, elicitation is separable from capabilities development. Note, though, that in practice, the two are also closely tied. That is, when we speak about the various incentives in the world that push towards capabilities development, they specifically push towards the development of capabilities that you are able to elicit in the way you want. If the capabilities in question remain locked up inside the model, thatâs little help to anyone, even the most incautious AI actors who are âfocusing solely on capabilities.â
Admittedly, itâs a little bit conceptually fuzzy what it takes for a capability to be âinâ a model, but for you to be unable to elicit it.
Here, weâre specifically talking about eliciting desired task-performance of a superintelligent agent that satisfies the agential pre-requisites and goal-content pre-requisites I describe here. So itâs natural, in that context, to use the agency-loaded frame in particular â that is, to talk about how the AI would evaluate different plans that involve using its capabilities in different ways.[18]Â
And if weâre thinking in these terms, we can modify the framework I used re: takeover seeking above to reflect an important difference between various non-takeover options: namely, that some of them involve doing the task in the desired way, and some of them do not. In a diagram:
That is: above we discussed our prospects for avoiding a scenario where the AI chooses its favorite takeover option. But in order to get desired elicitation, we need to do something else: namely, we need to make sure that from among the AIâs non-takeover options, it specifically chooses to âdo the task in the desired way,â rather than to do something else.[19] (Letâs assume that the AI knows that doing the task in the desired way is one of its options â or at least, that trying to do the task in this way is one of its options.)
Ok, those were some comments on desired elicitation. Now I want to say a few things about the role of âverificationâ in the dynamics discussed so far.
In my discussion of the âverificationâ in section 2, I said above that we donât, strictly, need to âverifyâ that our aims with respect to ensuring safety properties (i.e., avoiding takeover) or elicitation properties are satisfied with respect to a given AI â what matters is that they are in fact satisfied, even if we arenât confident that this is the case. Still, I think verification plays an important role, both with respect to avoiding takeover, and with respect to desired elicitation â and I want to talk about it a bit here.
Here Iâm going to use the notion of âverificationâ in a somewhat non-standard way, and say that you have âverifiedâ the presence of some property X if you have reached justifiably levels of confidence in this property obtaining. This means that, for example, youâre in a position to âverifyâ that there isnât a giant pot of green spaghetti floating on the far side of the sun right now, even though you havenât, like, gone to check. This break from standard usage isnât ideal, but Iâm sticking with it for now. In particular: I think that ultimately, âjustifiable confidenceâ is the thing we typically care about in the context of verification.
Letâs say that if you are proceeding with an approach to the alignment problem that involves not verifying (i.e., not being justifiably confident) that a given sort of property obtains, then you are using a âcross-your-fingersâ strategy.[20]Â Such strategies are indeed available in principle. And I suspect that they will be unfortunately common in practice as well. But verification still matters, for a number of reasons.
The first is the obvious fact that cross-your-fingers strategies seem scary. In particular, insofar as a given type of safety property is critical to avoiding takeover/omnicide (e.g., a property like âwill not try to takeover on the input Iâm about to give itâ), then ongoing uncertainty about whether it obtains corresponds to ongoing ex ante uncertainty about whether youâre headed towards takeover/omnicide.
Even absent these âwe all die if X property doesnât obtainâ type cases, though, it can still be very useful and important to know if X obtains, including in the context of capability-elicitation absent takeover. Thus, for example, if we want our superintelligent AI agent to be helping us cure cancer, or design some new type of solar cell, or to make on-the-fly decisions during some kind of military engagement, itâs at least nice to feel confident that itâs actually doing so in the way we want (even if weâre independently confident that it isnât trying to take over).
Whatâs more: our ability to verify that some property holds of an AIâs output or behavior is often, plausibly, quite important to our ability to cause the AI to produce output/behavior with the property in question. That is: verification is often closely tied to elicitation. This is plausible in the context of contemporary machine learning, for example, where training signals are our central means of shaping the behavior of our AIs. But it also holds in the context of designing functional artifacts more generally. I.e., the process of trying something out, seeing if it has a desired property, then iterating until it does, will likely be key to less ML-ish AI development pathways too â but the âseeing if it has a desired propertyâ aspect requires a kind of verification.
Letâs look at our options for verification in a bit more depth.
Suppose that you have some process P that produces some output O. In this context, in particular, weâre wondering about a process P that includes (a) some process for creating a superintelligent AI agent, and (b) that AI agent producing some output â e.g., a new solar cell, a set of instructions for a wet-lab doing experiments on nano-technology, some code to be used in a companyâs code-base, some research on alignment, etc.
Youâd like to verify (i.e., become justifiably confident) that this output has some property X â for example, that the solar cell/wet-lab/code will work as intended, that it wonât lead to or promote a takeover somehow, etc. What would it take to do this?
We can distinguish, roughly, between two possible focal points of your justification: namely, output O, and process P. Letâs say that your justification is âoutput-focusedâ if it focuses on the former, and âprocess-focusedâ if it focuses on the latter.
Most real-world justificatory practices, re: the desirability of some output, mix output-focused and process-focused justification together. Indeed, in theory, it can be somewhat hard to find a case of pure output-focused justification â i.e., justification that holds in equal force totally regardless of the process producing the output being examined.
One candidate purely output-focused justification might be: if you ask any process to give you the prime factors of some semiprime i, then no matter what that process is, youâll be able to verify, at least, that the numbers produced, when multiplied together, do in fact equal i (for some set of reasonable numbers, at least).[21]Â
E.g., at least within reasonable constraints, even a wildly intelligent superintelligence canât give you two (reasonable) numbers, here, such that youâll get this wrong.[22]
Indeed, in some sense, we can view a decent portion of the alignment problem as arising from having to deal with output produced by a wider and more sophisticated range of processes than weâre used to, such that our usual balance between output-focus and process-focus in verifying stuff is disrupted. In particular: as these processes are more able to deceive you, manipulate you, tamper with your measurements, etc â and/or as they are operating in domains and at speeds that you canât realistically understand or track â then your verification processes have to rely less and less on sort of output-focused justification of the form âI checked it myself,â and they need to fall back more and more either on (a) process-focused justification, or (b) on deference to some other non-correlated process that is evaluating the output in question. Â
Correspondingly, I think, we can view a decent portion of our task, with respect to the alignment problem, as accomplishing the right form of âepistemic bootstrapping.â[23]Â That is, we currently have some ability to evaluate different types of outputs directly, and we have some set of epistemic processes in the world that we trust to different degrees. As we incorporate more and more AI labor into our epistemic toolkit, we need to find a way to build up justifiable trust in the output of this labor, so that it can then itself enter into our epistemic processes in a way that preserves and extends our epistemic grip on the world. If we can do this in the right order, then the reach of our justified trust can extend further and further, such that we can remain confident in the desirability of whatâs going on with the various processes shaping our world, even as they become increasingly âbeyond our kenâ in some more direct sense.
Now, above I mentioned a general connection between verification and elicitation, on which being able to tell whether youâre getting output with property X (whether by examining the output itself, or by examining the process that created it) is important to being able to create output with property X. In the context of ML, we can also consider a more specific hypothesis, which I discussed in my post âThe âno sandbagging on checkable tasksâ hypothesis,â according to which, roughly, the ability to verify (or perhaps: to verify in some suitably output-focused way?) the presence of some property X in some output O implies, in most relevant cases, the ability to elicit output with property X from an AI capable of producing it.
In that post, I didnât dwell too much on what it takes for something to be âcheckable.â The paradigm notion of âcheckability,â though, is heavily output-focused. That is, roughly, we imagine some process that mostly treats the AI as a black box, but which examines the AIâs output for whether it has the desired property, then rewards/updates the model based on this assessment. And the question is whether this broad sort of training would be enough for desired elicitation.
If the âno sandbagging on checkable tasksâ hypothesis were true of superintelligent AI agents, for a heavily output-focused notion of checkable, and you could make the task performance you want to elicit output-focused-âcheckableâ in the relevant sense, then you could get desired elicitation this way. And note, as ever, that the type of output-focused checkability at stake, here, can draw on much more than unaided human labor. That is, we should imagine humans assisted by AIs doing whatever we justifiably trust them to do (assuming this trust is suitably independent from our trust in the process whose output is being evaluated). This is closely related to our prospects for âscalable oversight.â
In general, I think itâs an interesting question exactly how difficult it would be to output-verify the sorts of task-performance at stake in âaccess to the main benefits of superintelligent AI.â For various salient tasks â e.g. curing cancer, vastly improving our scientific understanding, creating radical abundance, etc (I think it would be useful to develop a longer list here and look at it in more detail) â my suspicion is that we can, in fact, output-focused verify much of what we want, at least according to the normal sorts of standards we would use in other contexts. E.g., and especially with AI help, I think we can probably recognize a functional and not-catastrophically-harmful cancer cure, solar cell, etc if our AIs produced one.
However, at the least, and even in the context of heavily output-focused forms of âchecking,â I think we are likely going to need some aspect of process-focused verification as well, to rule out cases where the AIs are messing with our output-focused verification in more sophisticated ways â e.g., faking data, messing with measurement devices, etc.[24]
More broadly, though, it also seems possible that even if we can rule out various flagrant forms of measurement tampering, much of the task-performance we want out of superintelligent agents will end up quite difficult to verify in an output-focused way, even using scalable methods. For example, maybe this task performance involves working in a qualitatively new domain that even our scalable-oversight methods canât âreachâ epistemically.
Given the possible difficulties with relying centrally on output-focused verification, what are our options for more process-focused types of verification?
I wonât examine the issue in much depth here, but here are a few routes that are currently salient to me:
Imitation learning: another sort of process-focused argument you could give would be something like: âwe trained this agent via imitation learning on human data to be like a human in a blah way. We claim that in virtue of this, we can trust it to be producing output with property X in blah context we canât output-verify.â[25]Â
Plausible that this is actually just a sub-variant of a âgeneralization + 'no successful adversariality'â arg. That is, plausibly you need to really be saying âit was like a human in blah way in these other contexts, and if it remains like a human in blah way in this context we canât output-verify than things are good, and we do expect it to generalize in this way for blah reasons (including: that itâs not being successfully actively adversarial).â But I thought Iâd flag it separately regardless.
A few other notes:
In general, I expect our actual practices of verification to mix output-focus and process-focus together heavily. E.g., you try your best to evaluate the output directly, and you also try your best to understand the trustworthiness of the process â and you hope that these two, together, can add up to justified confidence in the outputâs desirability.
I want to close with a discussion of whether solving the alignment problem in the sense Iâve described requires some very sophisticated philosophical (not to mention technical) achievement â and in particular, whether it requires successfully pointing an AI at some object like our âvalues on reflection,â our âcoherent extrapolated volition,â or some such.
As I noted above, I think the alignment discourse is haunted by some sense that this sort of philosophical achievement is necessary.
My current guess, though, is that we donât actually need to successfully point at (and get an AI to care intrinsically about) some esoteric object like our âvalues on reflectionâ in order to solve alignment in the sense Iâve outlined. And good thing, too, because I think our âvalues on reflectionâ may not be a well-defined object at all.
One intuition pump here is: in the current, everyday world, basically no one goes around with much of a sense of what peopleâs âvalues on reflectionâ are, or where they lead. Rather, we behave in desirable ways, vis-a-vis each other, by adhering to various shared, common-sense norms and standards of behavior, and in particular, by avoiding forms of behavior that would be flagrantly undesirable according to this current concrete person â or perhaps, according to some minimally extrapolated version of this person (i.e., what this person would think if they knew a bit more about the situation, rather than about what they would think if they had a brain the size of a galaxy).
Whatâs more, and even if we do end up needing to deal with edge cases or with a bunch of gnarly ethical/philosophical questions in order to get non-takeover/desired elicitation from our AIs, I think itâs plausible that getting access to something like an âhonest oracleâ â that is, an AI that will answer questions for us honestly, to the best of its ability â is enough to get us most of what we want here â and indeed, perhaps most of whatâs available even in principle. And I think an âhonest oracleâ is a meaningfully more minimal standard than âan AI that cares intrinsically about your values-on-reflection.â
Here Iâm roughly imagining something like: if you have an honest oracle, you can in principle ask it a zillion questions like: âif we do blah thing, is it going to lead to something I would immediately regret if I knew about it,â âwhat would I think about this thing if ten copies of me debated about it in the following scenario for the following amount of time,â âis there something about this thing that Iâd probably really want to know that I donât know right now?,â etc.[26]Â And as I discussed in âon the limits to idealized values,â I think the full set of answers to questions like this is probably ~all that the notion of your âvalues on reflectionâ comes down to.
That is, ultimately, there is just the empirical pattern of: what you would think/feel/value given a zillion different hypothetical processes; what you would think/feel/value about those processes given a zillion different other hypothetical processes; and so on. And you need to choose, now, in your actual concrete circumstance, which of those hypotheticals to give authority to.
So in a sense, on this picture, an honest oracle would give you access to ~everything there is to access about your values on reflection. The rest is on you, now.
Now, of course, there are lots of questions we can raise about ways that honest oracles can be dangerous, and/or extremely difficult, in themselves, to create (though note that an honest oracle doesnât need to be a unitary mind â rather, it just needs to be some reliable process for eliciting the answers to the questions at stake). And as I noted above, notions like honesty, non-manipulation, and so on do themselves admit of various tough edge cases. Iâm skeptical, though, that resolving all of these edges adequately itself requires reference to our full values-on-reflection (i.e., I think that good-enough concepts of âhonestyâ and ânon-manipulationâ are likely to be simpler and more natural objects than the full details of our full-values-on-reflection, whatever those are). And as above, I think itâs plausible that if you can just get AIs that arenât dishonest or manipulative in non-edge-case ways, this goes a ton of the way.
We can also ask questions about how far we could get with more minimal sorts of âoracleâ-like AIs. Thus, an âhonest oracleâ is intuitively up for trying to answer questions about weird counterfactual universes, somewhat ill-specified questions, and the like â questions like âwould I regret this if a million copies of me went off into a separate realm and thought about it in blah way.â But we can also consider âprediction oraclesâ that only answer questions about different physically-possible branches of our current universe, âspecified-questionâ oracles that only answer questions specified with suitable precision, and the like. And these may be easier to train in various ways.[27]Â
OK, those were some disparate reflections on whatâs involved in solving the alignment problem. Admittedly, itâs a lot of taxonomizing, defining-things, etc â and itâs not clear exactly what role this sort of conceptual work does in orienting us towards the problem. But Iâve found that for me, at least, itâs useful to have a clear picture of what the high level aim is and is not, here, so that I can keep a consistent grip on how hard to expect the problem to be, and on what paths might be available for solving it.
This is a somewhat deviant definition, in that it doesnât require that youâve created a superintelligence that is in some sense aimed at your values/intentions etc. But thatâs on purpose.
The term "epistemic bootstrapping" is from Carl Shulman.
I have to specify âbad,â here, because some conceptions of alignment that Iâll discuss below countenance âgoodâ forms of AI takeover.
And more generally, it seems like to me that ensuring that humanity gets the benefits of as-intelligent-as-physically-possible AI, even conditional on getting the benefits of superintelligence, is very much not my job.
Thanks to Ryan Greenblatt for conversation on this front.
Thanks to Ryan Greenblatt for discussion.
This is going to be relative to some development pathway for those more capable models.
Iâll count it as âuncoordinatedâ if many disparate AI systems go rogue and succeed at escaping human control, but then after fighting amongst themselves one faction emerges victorious.
In principle different AI systems participating in a coordinated takeover could predict different odds of success, but Iâll ignore this for now.
If misaligned AIs end up controlling ~all future resources, but humans end up with some tiny portion, Iâll say that this still counts as a takeover â albeit, one that some human value systems might be comparatively OK with.
I grant that a sufficiently superintelligent agent would have a DSA of this kind; but whether the least-smart agent that still qualifies as âsuperintelligentâ would have such an advantage is a different question.
 I focus on actions directly aimed at takeover here, but to the extent that uncoordinated takeovers involve AIs acting to secure other forms of more limited power, without aiming directly at takeover, a roughly similar analysis would apply â i.e., just replace âtakeoverâ with âsecuring blah kind of more limited powerâ; and of think of âeasinessâ in terms of how easy or hard it would be for the effort to secure this power to succeed.
 See Lethality 2: âA cognitive system with sufficiently high cognitive powers, given any medium-bandwidth channel of causal influence, will not find it difficult to bootstrap to overpowering capabilities independent of human infrastructure.â Though note that âsufficiently highâ is doing a lot of work in the plausibility of this claim â and our real-world task need not necessarily involve building an AI system with cognitive powers that are that high.
 Here I think we should be interpreting the input in question in terms of the sorts of âcorrectionsâ at stake in Yudkowskyâs notion of âcorrigibilityâ â e.g., shutting down the AI, or changing its values. A benign sovereign AI might still give humans other kinds of input â e.g., because it might value human autonomy (though I think the line between this and âcorrigibilityâ might get blurry).
 And note that to meet my definition of âsolving the alignment problem without access-to-benefits loss,â weâd need to assume that âsomehow OKâ here means that those benefits are relevantly accessible.
 Of course, depending on the specific way it obeys instructions, you can potentially turn a loyal assistant into something like an âagent that shares your valuesâ by asking it to just act like an agent that shares your values and to ignore all future instructions to the contrary. But the two categories remain distinct.
I then have to modulate this vision to accommodate concerns about moral patienthood.
Note, though, that this approach brings in a substantive assumption: namely, that to the extent you are eliciting desired task-performance from the AI in question, you are specifically doing so from the AI qua potentially-dangerous-agent. That is, when the AI is doing the task, it is doing so in a manner driven by its planning capability, employing its situational awareness, etc.
Itâs conceptually possible that you could get desired task performance without drawing on the AIâs dangerous agential-ness in this way. E.g., the image would be something like: sure, sometimes the AI sits around deciding between take-over plans and other alternatives, and having its behavior coherently driven by that decision-making. But when itâs doing the sorts of tasks you want it to do, itâs doing those in some manner that is more on âautopilot,â or more driven by sphex-ish heuristics/unplanned impulses etc.
That said, this approach starts to look a lot like âbuild a dangerous SI agent but donât use it to get the benefits of superintelligence.â E.g., here youâve built a dangerous SI agent, but youâre not using it qua dangerous to get the benefits of superintelligence. At which point: why did you build it at all?
Because this is specifically an elicitation problem, weâre assuming that the AI has this as an option.
Obviously, in reality there are different degrees of crossing-your-fingers, corresponding to different amounts of justifiable confidence, but letâs use a simple binary for now.
Iâm setting aside whether you can verify that those numbers are prime.
Note that youâre allowed to use tools like calculators here, even though your reasons for trusting those tools might be âprocess-inclusive.â What matters is that your justification for believing that property X holds makes minimal reference to the process that produced the output in question, or to other processes whose trustworthiness is highly correlated with that process (the calculatorâs trustworthiness isnât).
This is a term from Carl Shulman.
Thanks to Ryan Greenblatt for extensive discussion here.
Thanks to Collin Burns for discussion.
Thanks to Carl Shulman and Lukas Finnveden for discussion here.
See e.g. the ELK reportâs discussion of ânarrow elicitation,â and the corresponding attempt to define a utility function given success at narrow elicitation, for some efforts in this vein (my impression is that an âhonest oracleâ in my sense is more akin to what the ELK report calls âambitious ELKâ â though maybe even ambitious ELK is limited to questions about our universe?).
Executive summary: Solving the AI alignment problem involves building superintelligent AI agents, avoiding bad forms of AI takeover, gaining access to the main benefits of superintelligence, and being able to elicit some of those benefits from the AI agents, without necessarily requiring the AIs to have human-aligned values or goals.
Key points:
Â
Â
This comment was auto-generated by the EA Forum Team. Feel free to point out issues with this summary by replying to the comment, and contact us if you have feedback.