I'm not sure this is as obvious as you're making out. If you give a misaligned ASI a goal to maximize its own wellbeing it could still do that very well, it may also just do other things you didn't intend for it to do (like wipe out humans).
A paperclip maximizer does its job really well after all, and we wouldn't call it aligned.
I think the article gives the impression that EAs are generally happy for humans to go extinct if they are replaced by happy AIs. But this isn't my impression. EA, at least outwardly, is very concerned about human extinction! Almost obsessively so!
I think a utilitarian can be in favor of humans staying in control and surviving. They might not think AIs will be capable of wellbeing, or might think there is a good chance they will have negative wellbeing, or might not be confident they will create a flourishing future.
What annoyed me most about the Hendryck's article is that he essentially assumes impartial utilitarianism must be wrong because it can lead to a conclusion he (and others) don't like. Impartial utilitarianism also would have suggested to slave owners they shouldn't keep slaves - a conclusion they and others wouldn't have liked!
I would have liked more philosophical discussion about if impartiality is or is not a good moral principle. But I suppose he wasn't interested in getting into that.
This is cool! Did you consider testing for the very repugnant conclusion? Maybe I'm misremembering but I don't think any of your choices tested if we can just outweigh adding suffering people by adding more happy people.
A population of arbitrarily many lives with arbitrarily high welfare is worse than a population of arbitrarily many arbitrarily negative lives plus sufficiently many “ε-lives”[1] that each have an arbitrarily small quantity of positive welfare (Figure 4.1).
Forethought's view that improving the future conditional on survival is more important than ensuring survival goes against the dominant view in EA for many years that we need to reduce extinction risk. Two questions on this:
How far away from the optimal allocation of (longtermist) resources do you think the community currently is?
For example, should we be radically reducing investment in things like addressing biorisk or nuclear risk? Do we need to be rethinking the allocation of resources within AI risk?
Do you think there is anything that is being prioritized in the community that is actually harmful?
For example, could certain AI alignment approaches be bad for future digital sentience?
I haven't yet got past the 1.4. Arguments for Value of Variety section because I'm just a bit unconvinced.
You could reword the intuition pump section like this:
Imagine some truly terrible moment — say extreme torture. Suppose that this torture is far more terrible than anything that humanity has experienced to date: you or I would give up years of ordinary happy life just to avoid such a peak of despair. But now suppose that this torture is just ever so slightly less bad than some other torture approach that is the worst thing that could conceivably be produced, with the same resources. For example, a radically different device is used which leads to an experience that is ever so slightly more painful.
What is better? The worst possible torture alternating with the slightly less bad torture? Or just the slightly less bad torture for the rest of time?
I suppose I can imagine someone saying the former, but I wouldn't. I just want less suffering! You can dismiss this rewrite by saying variety is only good if it's variety of good things, but this would introduce an asymmetry and I'm unsure that is justified. I feel like people say they like variety because we have repeatedly experienced it to be pleasurable, and that introduces a bias that we struggle to avoid when we are asked to judge scenarios that aren't different in terms of welfare. For the same reason I'm a little unconvinced by the intrapersonal variety argument.
On the realisation-value argument. I don't really think there is intrinsic value of things being realized. If in the distant arctic some polar bear walks a route that no polar bear has walked before but which is the exact same in every welfare-relevant way, I just don't really care. Which is another way of me saying, realizing new things can indeed be great, but only when we can enjoy them for being new.
On the benefits for axiology point. This doesn't so much seem an argument for variety as it seems a direct argument for the saturation view. If the saturation view allows us to avoid lots of other unpalatable conclusions then it may be worth adopting for that alone!
cluelessness about some effects (like those in the far future) doesn’t override the obligations given to us by the benefits we’re not clueless about, such as the immediate benefits of our donations to the global poor
What makes you think that? Are you embracing a non-consequentialist or non-impartial view to come to that conclusion? Or do you think it's justified under impartial consequentialism?
cluelessness about some effects (like those in the far future) doesn’t override the obligations given to us by the benefits we’re not clueless about, such as the immediate benefits of our donations to the global poor
I do reject this thinking because it seems to imply either:
Embracing non-consequentialist views: I don't have zero credence in deontology or virtue ethics, but to just ignore far future effects I feel I would have to have very low credence in consequentialism, given the expected vastness of the future.
Rejecting impartiality: For example, saying that effects closer in time are inherently worth more than those farther away. For me, utility is utility regardless of who enjoys it or when.
The background assumption in this post is that there are no such interventions.
There's certainly a lot of stuff out there I still need to read (thanks for sharing the resources), but I tend to agree with Hilary Greaves that the way to avoid cluelessness is to target interventions whose intended long-run impact dominates plausible unintended effects.
For example, I don't think I am clueless about the value of spreading concern for digital sentience (in a thoughtful way). The intended effect is to materially reduce the probability of vast future suffering in scenarios that I assign non-trivial probability. Plausible negative effects, for example people feeling preached to about something they see as stupid leading to an even worse outcome, seem like they can be mitigated / just don't compete overall with the possibility that we would be alerting society to a potentially devastating moral catastrophe. I'm not saying I'm certain it would go well (there is always ex-ante uncertainty), but I don't feel clueless about whether it's worth doing or not.
And if we are helplessly clueless about everything, then I honestly think the altruistic exercise is doomed and we should just go and enjoy ourselves.
I'm not sure this is as obvious as you're making out. If you give a misaligned ASI a goal to maximize its own wellbeing it could still do that very well, it may also just do other things you didn't intend for it to do (like wipe out humans).
A paperclip maximizer does its job really well after all, and we wouldn't call it aligned.
I think the article gives the impression that EAs are generally happy for humans to go extinct if they are replaced by happy AIs. But this isn't my impression. EA, at least outwardly, is very concerned about human extinction! Almost obsessively so!
I think a utilitarian can be in favor of humans staying in control and surviving. They might not think AIs will be capable of wellbeing, or might think there is a good chance they will have negative wellbeing, or might not be confident they will create a flourishing future.
What annoyed me most about the Hendryck's article is that he essentially assumes impartial utilitarianism must be wrong because it can lead to a conclusion he (and others) don't like. Impartial utilitarianism also would have suggested to slave owners they shouldn't keep slaves - a conclusion they and others wouldn't have liked!
I would have liked more philosophical discussion about if impartiality is or is not a good moral principle. But I suppose he wasn't interested in getting into that.
Yeah I think that people with suffering-focused views could accept RC but not VRC.
RC doesn't consider anyone with a negative life, but VRC considers arbitrarily many arbitrarily negative lives.
I think it would be an interesting thing to add if you ever do a v2.
This is cool! Did you consider testing for the very repugnant conclusion? Maybe I'm misremembering but I don't think any of your choices tested if we can just outweigh adding suffering people by adding more happy people.
The Very Repugnant Conclusion (taken from here)
A population of arbitrarily many lives with arbitrarily high welfare is worse than a population of arbitrarily many arbitrarily negative lives plus sufficiently many “ε-lives”[1] that each have an arbitrarily small quantity of positive welfare (Figure 4.1).
Why is utility monster reasoning obviously wrong? Surely a lot of utilitarians just bite the bullet there?
I can't seem to delete post drafts. Delete option does not show when I click the three buttons to the right of the draft on my user page.
Forethought's view that improving the future conditional on survival is more important than ensuring survival goes against the dominant view in EA for many years that we need to reduce extinction risk. Two questions on this:
I haven't yet got past the 1.4. Arguments for Value of Variety section because I'm just a bit unconvinced.
You could reword the intuition pump section like this:
Imagine some truly terrible moment — say extreme torture. Suppose that this torture is far more terrible than anything that humanity has experienced to date: you or I would give up years of ordinary happy life just to avoid such a peak of despair. But now suppose that this torture is just ever so slightly less bad than some other torture approach that is the worst thing that could conceivably be produced, with the same resources. For example, a radically different device is used which leads to an experience that is ever so slightly more painful.
What is better? The worst possible torture alternating with the slightly less bad torture? Or just the slightly less bad torture for the rest of time?
I suppose I can imagine someone saying the former, but I wouldn't. I just want less suffering! You can dismiss this rewrite by saying variety is only good if it's variety of good things, but this would introduce an asymmetry and I'm unsure that is justified. I feel like people say they like variety because we have repeatedly experienced it to be pleasurable, and that introduces a bias that we struggle to avoid when we are asked to judge scenarios that aren't different in terms of welfare. For the same reason I'm a little unconvinced by the intrapersonal variety argument.
On the realisation-value argument. I don't really think there is intrinsic value of things being realized. If in the distant arctic some polar bear walks a route that no polar bear has walked before but which is the exact same in every welfare-relevant way, I just don't really care. Which is another way of me saying, realizing new things can indeed be great, but only when we can enjoy them for being new.
On the benefits for axiology point. This doesn't so much seem an argument for variety as it seems a direct argument for the saturation view. If the saturation view allows us to avoid lots of other unpalatable conclusions then it may be worth adopting for that alone!
I'll read those. Can I ask regarding this:
What makes you think that? Are you embracing a non-consequentialist or non-impartial view to come to that conclusion? Or do you think it's justified under impartial consequentialism?
I do reject this thinking because it seems to imply either:
There's certainly a lot of stuff out there I still need to read (thanks for sharing the resources), but I tend to agree with Hilary Greaves that the way to avoid cluelessness is to target interventions whose intended long-run impact dominates plausible unintended effects.
For example, I don't think I am clueless about the value of spreading concern for digital sentience (in a thoughtful way). The intended effect is to materially reduce the probability of vast future suffering in scenarios that I assign non-trivial probability. Plausible negative effects, for example people feeling preached to about something they see as stupid leading to an even worse outcome, seem like they can be mitigated / just don't compete overall with the possibility that we would be alerting society to a potentially devastating moral catastrophe. I'm not saying I'm certain it would go well (there is always ex-ante uncertainty), but I don't feel clueless about whether it's worth doing or not.
And if we are helplessly clueless about everything, then I honestly think the altruistic exercise is doomed and we should just go and enjoy ourselves.