I’ve always been attracted to Effective Altruism’s core principles of impartiality, scope sensitivity, scout mindset, and recognition of tradeoffs. But recently I’ve had a niggling feeling that the EA longtermist community has lost sight of these. In particular that EA has implicitly been assuming a special status for humans — at odds with the impartiality principle.
Below I go through some ways I think we may be going wrong.
Mitigating the risks of human extinction has long been a priority of the EA movement. The basic idea is that, if we go extinct, we miss out on what could be a big, flourishing future. Worries about human extinction have at least partly driven several priorities including preventing loss of control of advanced AI, preventing catastrophic pandemics, and preventing great power conflicts.
But the overwhelming importance of preserving humanity isn’t obvious under an impartial view. Firstly, it isn’t clear to me that humanity continuing is positive in expectation. Human survival could mean spreading factory farming to the stars. It could mean creating digital beings that suffer. I’m highly uncertain, but if I had to guess, humanity has had a net negative impact to date due to the immense scale of suffering on factory farms. Overall, we should be at least skeptical humanity continuing will be good.
It is also plausible that life could evolve again after human extinction and that futures in which we avoid an extinction event may not be all that great. Also, some extinction events could kill all wild animals, who plausibly suffer more than they flourish.
These arguments don't necessarily make human survival bad in expectation, but they're enough to blunt the case that avoiding extinction is obviously overwhelmingly important. This raises the question if there are more impactful things we can do (there are—I’ll get to this later).
An emerging priority in EA is avoiding the disempowerment of humans — scenarios where humanity continue to exist but its interests are sidelined and control over the future undermined. This could happen from advanced AI.
But key arguments for human disempowerment being bad assume that humans would have created a great future, which I have argued is unclear. Or more precisely, that humans would have created a future better than the actor that disempowered them (most likely advanced AI). In a Forethought piece, Tom Davidson considers that human takeover might be worse than AI takeover. In short, AIs are currently much nicer than humans and will be more competent, so they may be better at managing a flourishing future. Of course this is a contentious argument, and just because AIs are nice in training doesn’t mean they will be nice post-takeover. But this should again blunt the desirability of avoiding human disempowerment.
That said, there may be reasons to avoid human disempowerment that aren’t human-centric. If advanced AI displaces humans from work, greater power and wealth could be concentrated in the hands of capital owners who would then have control over the future. Power concentration could lead to tyranny and missing out on great futures. This is a valid concern under impartiality, and doesn’t require any belief about human specialness. But it's a concern about which humans hold power, not about humans losing control to AI. It’s worth keeping those two worries separate so we target the mechanisms that matter most, rather than disempowerment in general.
The classic AI alignment problem is that we don’t know how to give AI systems a goal without them doing stuff we don’t want them to do. They may seek power and resources to achieve their goal. They may go to great lengths to avoid being turned off. They may escape secure sandboxes, communicate with other AIs, and then hack into a third party.
The goal of much alignment research has been to figure out how to get AIs to do what we want without these negative consequences. The problem with this is that, even if we succeed, bad actors can still use AI to do bad things. There’s also a lot to lose if we simply use AIs as tools to do our bidding as we are fallible—we often make moral and empirical errors. Getting AI right should mean having at our fingertips an extremely powerful tool that can not only help us achieve our goals, but advise when these goals might be misguided and refuse to do bad things. This requires much more than simple alignment.
Will MacAskill has suggested that AIs should be “value-aligned”. A value-aligned AI wants to do good stuff. This could start with the AI being motivated by what we (humanity) currently consider good values, but importantly such an AI should also have built in processes to reflect, update values, and guide us towards a flourishing future. I do think that such an AI should be transparent about its reasoning and motives, and still be somewhat corrigible and limited in what it can do. But simply aligning AIs to human values isn’t good enough.
A common thread through each of my examples is assuming some special status for humans that may not stand up to scrutiny when considering EAs core principles. The community assumes it is overwhelmingly important for humans to survive, remain in power, and set the direction of the future. But if we take impartiality seriously I am unsure this is true. This isn’t to say it wouldn’t be good for humans to remain empowered, just that it’s highly unclear this should be a top priority for impartial altruists.
We are at a truly pivotal moment with the advent of advanced AI, which could set us on irreversible paths. It’s time for us to embrace our scout mindset, check our biases, and do what we can to improve the value of the far future as effectively as possible.
It is natural to want our own species to maintain a privileged place in the world, and I still want humans to flourish, I just don't think that should be the top priority on the current margin. In particular, we should start taking digital welfare much more seriously.
There could be an enormous number of digital minds in the future, and it’s possible that decisions we make in the coming decades could have long-lasting effects for their welfare. We can’t be sure that digital minds will be capable of welfare, but expert surveys suggest a decent probability they might be. If we take impartiality and scope sensitivity seriously, it seems to me digital minds should be near the top, if not the top priority for EAs. The stakes are just too high.
To the community’s credit, digital minds do get attention, and it seems concern is growing. But I don’t think the attention is high enough. 80,000 Hours does not rate digital minds in their list of core problems because “we know of fewer high-impact opportunities to work on this issue than on our top priority problems”. If the stakes are so high and we don’t know about high impact opportunities then perhaps an overwhelming priority should be to find these opportunities urgently. While the digital minds space may currently be less able to absorb human talent than others, 80,000 Hours relegating it to an “emerging priority” implies it isn’t yet a priority. I think this is wrong.
There is also a lack of attention towards potential conflicts between AI safety and AI welfare. AI safety usually proceeds just taking into account impacts on current and future humans, but we shouldn’t lose sight of the potential impacts on digital beings as we develop and manipulate them to serve human interests.
We should also embrace a Better Futures view. Rather than focus on ensuring our survival, which I have argued is of uncertain value, we should look to improve the quality of futures conditional on our survival. This is much more robustly good. A simple argument for this approach is illustrated in the diagram below. If we are closer to solving survival than we are to solving flourishing, there is a lot more to gain from solving flourishing. I think Forethought’s Better Futures series argues convincingly that this is the case.
Importantly, there are tangible things we can do to improve the future, conditional on our survival. Will MacAskill summarizes various actions under three key approaches:
EA has been a phenomenally successful movement. It has raised the plight of factory-farmed and wild animals. It alerted the world to the risks of advanced AI. But EA may have drifted from its principles. EA should do what it does best—remain impartial, search for big impact, and not be afraid to change course.