This competition entry has been selected for publication by the Forum team.
Here are some comments on the Summary of the argument.
What would justify preferring action A over B on impartial altruistic grounds? We’d need to “expect” that according to our epistemically idealized self, A has better expected total consequences across the cosmos (normative premise).
I disagree with the decision being made with respect to an epistemically idealized self, rather than with the information and reasoning ability available to me. Consider: if I am assigning credences to a coin flip, perhaps an idealized self would be able to simulate a coin flip really granularly, and be able to predict it perfectly, or to run lots of simulations and say that heads (or tails) would be 55%. But with my meager ressources, I can just say 50%. Not being able to say what the idealized reasoning doesn't expect me from having some expectation about how things will turn out.
Let’s say that we c-prefer[1] A over B if the reason we prefer A is an impartial altruistic comparison of the actions’ possible consequences
Ok
P1. Normative premise: To justify c-preferring A over B, it’s not enough to say (e.g.) that A seems heuristically good. Rather, we need to argue that A has higher “expected value” broadly speaking, meaning: In some sense we “expect” that, if we were idealized agents who could aggregate all of A’s and B’s possible consequences into literal EVs, then we’d say A has higher EV. (We ourselves don’t need literal EVs to justify c-preferences, hence the scare quotes.[2])
Not so! Requiring that expected values are well defined is actually a pretty strong requirement, which is not strictly needed to compare two options. We can also retreat to statewise. stochastic dominance, or even weaker versions if we are can't compute the expected value; see this point written up here.
For instance, with my limited reasoning abilities, perhaps I can't compute, for instance, the expected probability of xrisk if I work to prevent it vs if I don't. I'd say its about 0.2% to 10% either way. And yet I could be confident enough that the worlds in which I do work to prevent it stochastically dominate the worlds in which I do.
For another more detailed example, suppose I am comparing taking action X or not taking it. If I take it, there some unknown probability p that things would improve by P, and some unknown probability q that they will become worse by an amount Q. I have reason to believe that q<p and Q<P, but I’m very uncertain about what the amounts (or distributions) would be. Then nowhere are X, or even p, q, P, Q defined, the expected value surely isn’t defined, and yet I prefer taking that action over not doing so.
Otherwise, it’s unacceptably arbitrary to c-prefer A.
No, it is not arbitrary, you can c-prefer A over B given the information and reasoning ability available to you, without reference to what an idealized agent would do.
If our understanding of A’s and B’s possible consequences is sufficiently coarse-grained, then we don’t have an argument for “expecting” our idealized self’s EV for A to be higher, lower, or equal to B’s.[3] So A’s and B’s “EVs” are incomparable. In particular:
And yet we can still compare outcomes across expected trajectories. If you had to choose between the green and the red worlds, in the absence of other information, you would choose the green one, and you can thus use the same principle to make choices today, and have beliefs about the green line being preferable in the future, even if you can't quantify them. What is happening in the background is that you are doing something like "even though the future trajectories seem uncertain and could both be wide, in the absence of further information I still expect the green to be, on average, higher at any particular point in time and thus preferable overall", and this is correct.
P3. Empirical[4] premise: Due to unawareness (at least), our understanding of any pair of actions’ possible consequences is indeed very coarse-grained — enough that the conclusion of (P2) follows (i.e., these actions’ “EVs” are incomparable). In particular, the actions’ “EVs” are too severely imprecise to compare them, regardless of whether we (a) formally model these “EVs” or (b) appeal to informal/heuristic arguments
And as a result of the considerations above, this doesn't follow; we can be uncertain about the specific probabilities, or EVs, of two given actions, while still being able to believe in statewise or stochastic dominance (or I guess stochastic EV-dominance). I can have beliefs about me being able to lift less than a really buff guy even if I am uncertain about how much each of us could lift, and similarly I can have beliefs about the aggregate effect of an action affecting the future positively, in expectation, relative to another action, even though at some point the butterfly effect kicks in and I stop being able to model the granular futures.
Do you reject the normative premise, because you think it counts as an impartial altruistic justification if we say “This action has good ‘expected’ consequences after bracketing the consequences we’re clueless about”? Then arguably you should prioritize neartermist causes.
I don't think that bracketing would be justified, since I believe you can still form beliefs about the preferability of trajectories.
Do you accept the normative premise but reject the conceptual one, implying that we should form “best guesses” about the balance of all cosmos-wide consequences? Then you should look for interventions that are best after accounting for as many “galaxy-brained” considerations as possible, rather than simply ignore those considerations. (It’s been argued that mainstream x-risk reduction meets that bar — e.g., Shulman; Adelstein; Carlsmith — but I think this should be spelled out a lot more carefully.)
This does sound really cool indeed.
Do you agree with the normative and conceptual premises, but think some cautious or “meta” interventions are justified without arbitrary calls about the considerations we’re unaware of? Like, say, saving resources until we’re in a clearer epistemic situation? Then you should do those interventions, rather than various other popular interventions whose justification does rely on arbitrary calls
I do like that the Patient Philanthropy Fund exists.
I expect newer or sharper critiques of the normative premise, and part (b) of the conceptual premise, to be most productive
Yep.
“What is the standard that non-idealized impartial altruists (should) use to judge which actions are rational? If it’s ‘approximating EV’, or ‘going with our best guess, even if not with literal precise EVs’, what exactly do these things mean, and what justifies them?”
This isn't answering quite the same question, but I think the live question for me is: what should I be doing in the world given the limited information I have about it, my own constrains about reasoning and time, and my actual values, which include but are not limited to scale-sensitive components. And the answer is going to depend both on the values and the position I occupy in the world.
Third, relatedly, unawareness probably has some implications for impartial altruists, even if we don’t think it makes us clueless.
There is a related problem here, where I can see that some effective altruist actions historically backfired (investment in OpenAI and Anthropic, investing in Democrats ahead of the 2024 election and then losing, etc.), or notice that Musk helped start OpenAI only to later regret what it became. And notice: yeah, some actors in the world are predictably not humble enough about the effects of their own actions, perhaps including myself, and therefore I should acquire wisdom, study history, consult my elders, or else risk acting suboptimally like an immature adolescent. This basically seems correct. But then if I take actions that will make me seek wisdom this is again with reference to me thinking that the trajectory of the future will in expectation be better, even if I can't calculate the degree to which it will be.
Finally, if nothing else, it seems epistemically virtuous to be clear about the reasons for our decisions. Sure, perhaps there’s no behavioral difference between “I’m working on AI risk because I’ve really weighed up all the possible consequences, and it doesn’t seem arbitrary to say this work is impartially good ‘in expectation’”, and “I have no clue if my idealized self would favor working on AI risk, but I’m doing it because no one has offered something better”. But I think if we’re honest with ourselves that our reasoning is the latter, we’ll have more open minds if and when “something better” comes along
Sure.
Even if we’re forced to choose something, this doesn’t tell us whether we have impartial altruistic reasons to choose A or B.
FWIW, my impression is that thinking in terms of impartial altruistic reasons rather than in terms of your actual values (which may contain scale-sensitive components) is a mistake that leads to preference falsification.
By itself, “this heuristic favors action A” doesn’t tell us why A is c-preferable. We need to say why we believe this heuristic tracks A’s expected consequences from our idealized self’s perspective
No, we don't need to make reference to an idealized self, we can make reference to what we expect using the information that is available to us.
Point: Some actions are obviously c-preferable to others (not just preferable all things considered). So we should reject any philosophical argument to the contrary (“one person’s modus ponens is another’s modus tollens”). (Chappell; Mogensen)
Counterpoint: C-preferability is (arguably) not something we can directly perceive. Rather, it is constituted by weighing up possible consequences. So the justification for our beliefs about c-preferability depends on the justification for our beliefs about the consequences.[13] And because the set of consequences we need to weigh up is extremely complex, we can’t trust our intuitions about the bottom-line verdict “the weight of consequences favors A”. (2.3; see also “How to not do decision theory backwards”.)
I also cannot directly perceive what Putin is thinking, and the way the Russian state makes decisions might be extremely complex, and yet I can have credences about it. The way I think about it could be wrong, and yet if I am living in Russia I might want to leave before a war starts. I might be wrong. Either way I can't get a better brain, so I'm going to have to suck it up and take actions with the mental resources available to me. I am pretty sympathetic to the initial point.
Under uncertainty, we can precisely specify the possible outcomes we’re making tradeoffs between. But under unawareness, we can’t. So, since our values as impartial altruists are defined over very fine-grained possible outcomes, precise EVs aren’t well-defined.
cf. "I can have beliefs about me being able to lift less than a really buff guy even if I am uncertain about how much each of us could lift".
Point: Our intuitions about which actions are c-preferable are at least slightly better than chance. That is,[17] they’re positively correlated with the ground truth of “what we’d c-prefer if we could explicitly aggregate all the possible consequences”. That’s enough to always be able to say which action is c-preferable. (Lewis[18])
Counterpoint: We don’t have direct evidence that our intuitions about c-preferability tend to track truth. (2.3.1.1.) So we need to weigh (i) the weak positive evidence from (e.g.) near-term forecasting research, against (ii) other considerations, namely: First, our intuitions about c-preferability might systematically track things other than the truth (e.g., sources of bias in the sample of hypotheses that occur to us). (3.2.1.) Second, we should also put some weight on explicit models, which need to account for an extremely complex set of consequences. (3.2.) The problem is that it’s ambiguous how to weigh up (i) and (ii).[19] (2.1, 2.4.)
The counterpoint is making a type error. Yes, if we think that A>B, we do so because we have evidence that A>B and thus that thinking that A>B is correlated with A>B. Our intuitions might be wrong in either direction, but we don't have another brain, so we are just going to have to do as well we can.
Point: Even if our impact is dominated by consequences we’re unaware of, we don’t know which direction they point. So, subjectively we should regard the negative and positive consequences we’re unaware of as canceling out in expectation. (MacAskill[20]; Soares[21])
Counterpoint: It doesn’t follow from “we don’t know the net direction of the consequences we’re unaware of” that we should regard the positives and negatives as precisely symmetric. One reason symmetry is implausible: If we become aware of a new possible consequence, this should update our beliefs about the others we’re unaware of, breaking the symmetry. (4.1.1.)
Yes, it should follow, because the expected value of a probability should be itself, and the expected value of the expected value should be itself. Look up martingale. If you are not expecting your probabilities or your expected values to be a martingale, then you haven't incorporated all the information yet into them.
I think you might be thinking of a setup that might be: you think the expected value of something is 3, and you think that there will be 5 additional reasons with a +1 and 5 more with a -1. Then after learning that reason #1 is a +1, this "breaks the symmetry". But your expected value would still be 3, because you'd expect 4 additional +1 and 5 -1s. Ok, so what's the problem here?
Conversely, if you were originally thinking that the EV of something was a 3, and you get a consideration with a -1 (and say including your future expectation about future considerations, you increase that to a -1.5). Now you are at 1.5. You break no symmetry. You were as surprised about that -1.5 as you would have been by a +1.5, otherwise you wouldn't have had your future expected value.
For me, the decision problem is how to act in the world given limited information, and some portion of my actual values being scale-sensitive. Thinking in reference to an idealized actor seems counterproductive.
Assigning EVs and probabilities is a skill, which takes practice and which can be improved. There is a level of coarseness in which you are indeed pretty confused about everything, and a level of granularity in which you can make exquisitely finegrained distinctions. This is trainable.
What do I think is going wrong? You are making too many references to an idealized actor that is altruistic (fine) and has infinite information and compute (not fine) and becoming confused as a result.
What is my general solution? My general solution is that you can prefer trajectories over the future after a positive action over trajectories without that action, without needing to assign specific EVs or probabilities to either side. Separately, I think that you sometimes can estimate the probabilities and expected values (e.g., probabilities of xrisk given some action), but I think that’s not necessary to argue against cluelessness.
DiGiovanni advocates for with respect to your current expectation of what your epistemically idealized self would believe. This is nothing more than the reflection principle. If you don't think your idealized self believes X, why would you believe X? That'd be a violation of this very consensual principle.
(Idk how central this actually is to your overall case, but I wanted to react to that.)