the same questions can be raised for any philosophical question whatsoever.
I assume you didn't mean to restrict this to philosophical questions only, since the disagreement between Anti-Richard and you on nuclear war is largely empirical.[1]
If you tell me that jumping off a cliff would probably mean death, there are no/few Anti-Richards out there who will say the opposite. And there's a very good reason for that: the corpses of such anti-Richards are stacked at the bottom of the cliff. Evolution selected against them and in favor of your "jump --> death" judgment. So we have a reason to trust you over Anti-Richard on this. (And we can make a weaker and less straightforward version of this argument for some philosophical questions such as "is suffering bad?".)
The same can't be said for your judgment that "nuclear war is bad considering all the consequences from now until the end of time, including in scenarios where humanity survives and those where aliens replace us." If there are few/no Anti-Richards contradicting you, it's presumably not because of evolutionary pressure on correct beliefs on the matter.[2] So we have no reason to trust you over Anti-Richard, here.
You need to actually look at the substantive claims being made and judge for yourself what seems most credible.
So you're asking me to make a judgment call on whether to trust your opaque judgment call or Anti-Richard's. But unless we find a reason to believe one of the two tracks the truth better than the other, I don't see how I/you can be justified in trusting one of you more than the other.
Alpine and Beach are two incommensurably good holidays, between which your idealized self has no preference, even when $20 is added to either option. Still, Hare argues, you plainly ought to prefer the prospect of a mystery bag containing a ticket + $20 over a mystery bag containing just the other ticket. [...] this all holds despite the fact that you know that once you find out which ticket is in which bag, you will consider the two options to be incommensurable.
Isn't this conflating idealized agent with agent with more information? The reflection principle I'm defending is only about the former.
Thanks for developing on this, Richard. I think we can now identify more fine-grained cruxes than we did in our previous discussion, a bit.
Philosophical cluelessness, by contrast, commits one to the much stronger claim that nobody can (realistically) form any reasonable judgments or expectations about this question, no matter how carefully they investigate it. That’s an extremely strong skeptical claim!
You're claiming that you have a truth-tracking expectation of whether doing action A rather than B has an overall positive impact, considering, impartially, all their possible effects on all sentient beings from now until the end of time. I don't see how that's any less strong than suspending judgment on this very question, especially before being given any substantive argument for why I should c-prefer any A over any B (which your post does not do). I agree with Jo's comment.[1]
in many cases we can reasonably—albeit tentatively—expect “our idealized self’s EV for A to be higher, lower, or equal to B’s,”
If you really mean in many cases and not always, then you're objecting to P3 and not P2, I think.
(i) one option can be better in prospect than another, even when you know that, given more information, you would (correctly) regard the two outcomes as incommensurable or on a par; and (ii) in such cases, it’s rational to pick the better prospect rather than deferring to the idealized perspective.
This violates the reflection principle. I think you're going to need to at least give an argument if you want us to give up on this. Why would you believe something your idealized self tells you should not be believed?
since we’ve no grounds for expecting the balance of unknown reasons to count against rather than for our currently-preferred action, the fact of cluelessness is normatively inert: it makes no difference to our reasons for action.
Two possibilities. Either that's a version of the old "canceling out" objection to cluelessness that says the summed EV of the unknowns is exactly 0, and you don't engage with the many compelling rebuttals that have been given thus far (see, e.g., this, this, this, that, and references therein).
Or that's some form of wager based on bracketing out clueless worldviews (see, e.g., DiGiovanni's metanormative bracketing), which has important problems you're not addressing.
Value Correlation: An action’s short-run/visible value is a positive predictor of its long-run/invisible value. Actions that look good on the visible margin are more likely to be good than bad in the long run.
You don't seem to be advocating for human extinction because of humans' current massive negative impact on numerous farmed animals, which I'd bet you believe dominates humans' current welfare-relevant impact on themselves.
So presumably, you think the value correlation thesis is only one consideration among many, and not a slam-dunk one. This is just one thing in the pile of conflicting considerations that may make DiGiovanni and others clueless. This does not help them.
Relatedly, here's a comment with links to texts that clarifies why we should arguably suspend judgment on impartial goodness but not be, e.g., Pyrrhonian skeptics, and that the arguments for the former (those DiGiovanni endorses, at least) are very different from the arguments for the latter, contrary to what you seem to assume.
“In animal welfare, I can sometimes know that I have a solid chance of reducing the suffering of some animals in the near-term, and that’s potentially worth doing as long as you’re, uh, careful about the risks of doing harm. And the bracketing people have said that this can be good, so there’s at least one element of justification.”
I love how humble and honest this is. I wish everyone were as clear about how vague their rationale is and how much they defer to other people.
Sure, here are examples (and I'll take reasons not explicitly mentioned in the post, so this hopefully adds some value): - Reason to think expansion (such that Loic-the-ant/Dolores should not exist): that's what smart moral weight researchers have assumed so far. - Asymmetric reason to think granularization (such that Loic-the-ant/Dolores could very well exist): some subtle weak localized pains might require a complex brain architecture.
Here are symmetric reasons to contrast: - We might be in a simulation where the simulators are observing what sentience evolution through expansion looks like. - We might be in a simulation where the simulators are observing what sentience evolution through granularization looks like.
It seems legit to assume these symmetric reasons cancel out (that's "simple cluelessness"). Not so for the assymetric ones.
feel free to insert "person-moment" or "point of view occupied by a person" where my post says "person"; you'll get person-moment-centered bottom-up bracketing, or personal-point-of-view-centered bottom-up bracketing, based on exactly parallel arguments.
I've actually been confused about why Kollin et al. (2025) and Clifton (2025) never even mentioned something like "person-moments" as potential "locations of value". After a quick look, it seems they've been considered in infinite ethics, and I didn't find any argument for why they're clearly not suited as serious candidates.
Maybe any way of individuating person-moments was considered too arbitrary, or there's some reason, that does not come to mind rn, why bracketing over those would fail to provide action guidance?
I don't think I agree that her welfare based answer is obviously 'no'. If I put myself in her shoes, I think I would choose not to be cut by the glass.
I think Impartial was assuming that she should be clueless about the effect on her long-term well-being for similar reasons why we should be clueless about the impact of our actions on the long-term future (premise which we assume if we're considering bracketing in the first place)[1]. You can always find many reasons why the nasty cut would actually have good overall consequences (e.g., teaches her not to walk barefoot and she avoids potential future far nastier cuts), and our best guess on what reasons for and against sum up to does not seem truth-tracking. I struggle to see how one could defend that we know the nasty cut will overall harm her while holding that we should suspend judgment on the sign of any action A vs B on the overall future.
I think the real potential problem, though, is what you raise in this other comment: why should the units be "lifelong persons" rather than, say, person-moments? (See also this related thread started by Jesse.) Arguably, the reason why you would choose not to be cut by the glass if you were her, is because you are bracketing in near-term you (who is clearly harmed) and bracketing out longer-term you (where you're clueless). And, indeed, I curently fail to see how this move is any less warranted than what Impartial proposes, and the two moves are incompatible/contradictory.
The virtues you want to promote surely impact x-risks in a way that is not determinately good or bad in expectation, and it seems indeterminate whether this x-risk effect dominates. (Related argument here.)
This is only one of the many reasons we could find why we should be clueless about "whether these virtues systematically led to bad long-term consequences".
Consequentialist cluelessness seems infectious to your solution.
This is briefly addressed by DiGiovanni here and there (his cluelessness X near-term AW post is also relevant), and I discuss some aspect of his point in the first ref a bit further here.
Education (your example) may benefit AI accelerationists as much as, or more than, AI safety researchers. It may make humans better (for better or for worse) at creating digital minds, colonizing space, spreading wild animal suffering, or intensifying the farming of small animals off-Earth. We can't just assume education is going to solve all the problems you're worried about, rather than worsen them.
Just to be clear, is this your answer to the above: sure, but my overall best guess is that these possible negative consequences are outweighed by the good ones (e.g., because of good human bias thing), and my best guess is truth-tracking, unlike in many situations discussed by DiGiovanni. There is something special about education that breaks the usual paralysis from unawareness.
So, the cases seem like they are being driven, not by the fact that they are realistic or unrealistic, but by the fact that they specify the long-term causal consequences.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don't know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don't need (full) specification but a situation where the uncertainty is manageable.
I assume you didn't mean to restrict this to philosophical questions only, since the disagreement between Anti-Richard and you on nuclear war is largely empirical.[1]
If you tell me that jumping off a cliff would probably mean death, there are no/few Anti-Richards out there who will say the opposite. And there's a very good reason for that: the corpses of such anti-Richards are stacked at the bottom of the cliff. Evolution selected against them and in favor of your "jump --> death" judgment. So we have a reason to trust you over Anti-Richard on this. (And we can make a weaker and less straightforward version of this argument for some philosophical questions such as "is suffering bad?".)
The same can't be said for your judgment that "nuclear war is bad considering all the consequences from now until the end of time, including in scenarios where humanity survives and those where aliens replace us." If there are few/no Anti-Richards contradicting you, it's presumably not because of evolutionary pressure on correct beliefs on the matter.[2] So we have no reason to trust you over Anti-Richard, here.
Not sure what you mean by "substantive claim," but so far, you've given me i) an incredulous stare on the nuclear war thing, and ii) a value correlation argument, which I've argued is unhelpful.
So you're asking me to make a judgment call on whether to trust your opaque judgment call or Anti-Richard's. But unless we find a reason to believe one of the two tracks the truth better than the other, I don't see how I/you can be justified in trusting one of you more than the other.
Isn't this conflating idealized agent with agent with more information? The reflection principle I'm defending is only about the former.
The cruxes are the likelihood of a Smokey Bear effect and of aliens with better values taking over.
There are so many more plausible explanations, such as the fact that nuclear war seems bad for reasons that have nothing to do with impartiality and overall consequences until the end of time, and we incorrectly generalize.
Thanks for developing on this, Richard. I think we can now identify more fine-grained cruxes than we did in our previous discussion, a bit.
You're claiming that you have a truth-tracking expectation of whether doing action A rather than B has an overall positive impact, considering, impartially, all their possible effects on all sentient beings from now until the end of time. I don't see how that's any less strong than suspending judgment on this very question, especially before being given any substantive argument for why I should c-prefer any A over any B (which your post does not do). I agree with Jo's comment.[1]
I'm sure you'll agree that this could overall decrease x-risks because humanity could survive and be more peaceful/resilient afterwards. Or that human extinction might actually be good because there may be aliens with better values who'd take over if we're not there.
Say an alternative version of you (let's call them Anti-Richard) comes to us and says "my best guess is that starting a nuclear war is impartially good" because of the above considerations, and gives you an incredulous stare for thinking otherwise. You both agree on what the cruxes are but just make different opaque judgment calls. Why should I trust you any more than Anti-Richard? Why should you trust you any more than Anti-Richard? Why should I trust any of you any more than a coin flip?
If you really mean in many cases and not always, then you're objecting to P3 and not P2, I think.
This violates the reflection principle. I think you're going to need to at least give an argument if you want us to give up on this. Why would you believe something your idealized self tells you should not be believed?
Two possibilities. Either that's a version of the old "canceling out" objection to cluelessness that says the summed EV of the unknowns is exactly 0, and you don't engage with the many compelling rebuttals that have been given thus far (see, e.g., this, this, this, that, and references therein).
Or that's some form of wager based on bracketing out clueless worldviews (see, e.g., DiGiovanni's metanormative bracketing), which has important problems you're not addressing.
You don't seem to be advocating for human extinction because of humans' current massive negative impact on numerous farmed animals, which I'd bet you believe dominates humans' current welfare-relevant impact on themselves.
So presumably, you think the value correlation thesis is only one consideration among many, and not a slam-dunk one. This is just one thing in the pile of conflicting considerations that may make DiGiovanni and others clueless. This does not help them.
Relatedly, here's a comment with links to texts that clarifies why we should arguably suspend judgment on impartial goodness but not be, e.g., Pyrrhonian skeptics, and that the arguments for the former (those DiGiovanni endorses, at least) are very different from the arguments for the latter, contrary to what you seem to assume.
I love how humble and honest this is. I wish everyone were as clear about how vague their rationale is and how much they defer to other people.
And, fwiw, this sounds pretty sensible to me!
Sure, here are examples (and I'll take reasons not explicitly mentioned in the post, so this hopefully adds some value):
- Reason to think expansion (such that Loic-the-ant/Dolores should not exist): that's what smart moral weight researchers have assumed so far.
- Asymmetric reason to think granularization (such that Loic-the-ant/Dolores could very well exist): some subtle weak localized pains might require a complex brain architecture.
Here are symmetric reasons to contrast:
- We might be in a simulation where the simulators are observing what sentience evolution through expansion looks like.
- We might be in a simulation where the simulators are observing what sentience evolution through granularization looks like.
It seems legit to assume these symmetric reasons cancel out (that's "simple cluelessness"). Not so for the assymetric ones.
I've actually been confused about why Kollin et al. (2025) and Clifton (2025) never even mentioned something like "person-moments" as potential "locations of value". After a quick look, it seems they've been considered in infinite ethics, and I didn't find any argument for why they're clearly not suited as serious candidates.
Maybe any way of individuating person-moments was considered too arbitrary, or there's some reason, that does not come to mind rn, why bracketing over those would fail to provide action guidance?
I think Impartial was assuming that she should be clueless about the effect on her long-term well-being for similar reasons why we should be clueless about the impact of our actions on the long-term future (premise which we assume if we're considering bracketing in the first place)[1]. You can always find many reasons why the nasty cut would actually have good overall consequences (e.g., teaches her not to walk barefoot and she avoids potential future far nastier cuts), and our best guess on what reasons for and against sum up to does not seem truth-tracking. I struggle to see how one could defend that we know the nasty cut will overall harm her while holding that we should suspend judgment on the sign of any action A vs B on the overall future.
I think the real potential problem, though, is what you raise in this other comment: why should the units be "lifelong persons" rather than, say, person-moments? (See also this related thread started by Jesse.) Arguably, the reason why you would choose not to be cut by the glass if you were her, is because you are bracketing in near-term you (who is clearly harmed) and bracketing out longer-term you (where you're clueless). And, indeed, I curently fail to see how this move is any less warranted than what Impartial proposes, and the two moves are incompatible/contradictory.
If I refuse to work with this premise, then I should just do what I think is overall good and don't need bracketing. Problem solved.
Thanks, I meant the first.
The virtues you want to promote surely impact x-risks in a way that is not determinately good or bad in expectation, and it seems indeterminate whether this x-risk effect dominates. (Related argument here.)
This is only one of the many reasons we could find why we should be clueless about "whether these virtues systematically led to bad long-term consequences".
Consequentialist cluelessness seems infectious to your solution.
This is briefly addressed by DiGiovanni here and there (his cluelessness X near-term AW post is also relevant), and I discuss some aspect of his point in the first ref a bit further here.
Just to be clear, is this your answer to the above: sure, but my overall best guess is that these possible negative consequences are outweighed by the good ones (e.g., because of good human bias thing), and my best guess is truth-tracking, unlike in many situations discussed by DiGiovanni. There is something special about education that breaks the usual paralysis from unawareness.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don't know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don't need (full) specification but a situation where the uncertainty is manageable.