This competition entry has been selected for publication by the Forum team.
This is an option 3 entry, a constructive response. I accept Anthony's argument, including its conclusion. I'm not challenging a premise and I'm not saying the conclusion fails to follow. What I want to do is look at one modelling choice inside UEV, and argue that it changes what an impartial altruist should take away and where they should put their effort.
Anthony DiGiovanni's sequence argues that our unawareness of large-scale consequences runs so deep that we can't justify preferring one action to another on impartial grounds. I'm granting the empirical claim. Our conception of the far future really is coarse, there really are outcomes we haven't conceived of at all, and I don't think the response to that is a more confident guess or mild sweetening. I agree with Anthony's sequence almost entirely.
What I want to look at is a modelling choice in how UEV gets built, and what it makes hard to see afterwards.
~
Post 2 motivates UEV like this: our value function doesn't hand us precise values for hypotheses, so a natural move is to let the evaluation of a strategy be an interval rather than a number. The width of that interval is how imprecise we are. That's the whole idea, and it's a good one.
One thing to hold onto, because it matters later. The interval isn't meant to say "the true value is somewhere in here and I haven't found it yet." DiGiovanni is explicit that it reflects irreducible indeterminacy: nothing in our evidence or epistemic principles pins down a single value. That's a stronger and stranger claim than ordinary uncertainty, and I'm taking it at face value throughout.
Post 3's Appendix B then makes it exact. Roughly:
Carve up everything we're aware of into hypotheses. In the AI example: misaligned AI takes over, benevolent AI, malevolent human control, and a catch-all for what we haven't thought of.
Inside each hypothesis, imagine a set of much more specific worlds. These are the fine-grainings we only vaguely conceive of. Not "misaligned AI takes over" but each particular, fully-specified way that could go.
Define the value function over those specific worlds.
Let P be a set of probability distributions over which specific world we end up in.
The UEV of a strategy is the set of expected values you get, one for each distribution in P.
Here is the part worth pausing on. In that construction, the value function is a single fixed function. It doesn't vary. Every bit of imprecision in the final interval comes from P, the set of probability distributions.
So the thing that motivated the whole apparatus, that we can't precisely evaluate the hypotheses we can conceive of, doesn't end up represented as imprecision about value. It ends up represented as imprecision about which specific world obtains.
~
Not knowing what a hypothesis is worth, and knowing what each fine-grained version is worth but not which version you'll get, are close to interchangeable ways of describing the same epistemic situation. "I can't say what 'misaligned AI takes over' is worth" and "I know what each specific takeover scenario is worth, but I can't pin down which one we'd get" are two descriptions of one predicament. The second is more convenient to work with, so that's the one the formalism uses. Nothing has been smuggled in.
DiGiovanni also notes, in appendix B, that nothing in his argument turns on whether our credence in each hypothesis is imprecise. We might have perfectly precise credences at that level and still get his conclusion from imprecision over the fine-grainings inside each hypothesis. This appendix note is particularly important and I will return to its fine-grainings. I do think his modelling choices make sense and are defensible. My question now is about what it leaves out and how much it matters.
~
Post 2, section 2.2, gives two conditions under which unawareness stops being a big deal. A deep understanding of the mechanisms that determine a strategy's consequences. And inductive evidence of consistent success in similar contexts.
Both of those are questions about origin. Not "how wide is my interval" but "why is it wide, and could anything narrow it." Mechanistic understanding is something you can build by studying how a system works. A track record is something that accumulates when the world keeps checking your answers and you keep score.
Now look at what those criteria have to work with once the formalism is in place. There's one object, P, a set of probability distributions. The imprecision in it is a single quantity. You can ask how wide it is. You cannot ask which part of it came from not understanding a mechanism and which part came from not knowing what a world is worth, because the formalism no longer records that. It records extent, not origin.
The two criteria in 2.2 are origin-sensitive criteria being applied to an object that no longer carries origin information.
This matters because the two conditions don't fare the same across those origins. Take an imprecision that exists because we don't understand how AI development responds to a particular intervention. That is exactly what mechanistic study addresses, and exactly the sort of thing a forecasting track record could eventually discipline. Now take an imprecision that exists because we can't say what an unprecedented civilisational configuration is worth. Notice that both of 2.2's conditions are empirical. One is about understanding how a system works. The other is about how similar efforts have gone before. Neither is the kind of thing that speaks to an evaluative question at all. I'm not saying we lack the data yet. I'm saying that studying mechanisms harder, or gathering more cases of past success, isn't the sort of activity that bears on what an unprecedented world is worth. The two conditions aren't failing here for want of effort. They're aimed at a different kind of question.
I don't think DiGiovanni disagrees with the substance here. He's explicit in 3.3.1 that we're unaware at the level of both how good world-states are and how effective interventions are at steering between them, and his outcome robustness and implementation robustness table makes the point clearly. The distinction is in his prose, drawn deliberately. My claim is narrower: after Appendix B, there's no way to ask which side of that distinction a given interval's width came from, and 2.2's criteria are exactly the sort of criteria that would want to ask.
~
Return to the line about hypothesis-level credences. DiGiovanni says his argument doesn't need these credences to be imprecise. It runs fine on imprecision over the fine-grainings inside each hypothesis.
Whether an intervention shifts probability from one hypothesis toward another is about as close as this domain gets to a mechanistic, in-principle-checkable question. It's the kind of thing 2.2's criteria are built for. Not resolved today, but the sort of thing evidence could speak to.
Imprecision over the fine-grainings within a hypothesis is a different animal. Those are the specific unprecedented worlds we only vaguely conceive of, and the question of what they're worth is not one any mechanism study or track record is going to settle.
So the appendix note is telling us that the load-bearing imprecision, the part the conclusion actually rests on, sits precisely where 2.2's two conditions have least purchase. Appendix D does the same thing again, incidentally. All six of the standard approaches he considers, symmetry, extrapolation, meta-extrapolation, simple heuristics, focus on lock-in, capacity-building, get stated as constraints on EV*_p. So even the alternatives are expressed through the same probability object. I read that as the same pattern rather than a separate problem, but it does suggest the choice runs deeper than one appendix.
I don't read any of this as a flaw in the argument. If anything it makes the argument more robust, since it doesn't depend on the more tractable kind of uncertainty staying stuck. But it does mean the conclusion is better described than it currently is. Not "we're too unaware to compare actions," full stop, but "the imprecision that no evidence could ever narrow is by itself enough to block comparison."
This doesn't restore comparability. If the fine-grained imprecision alone does the work, then it does the work, and Maximality returns what it returns. Yet what should a reader take away from this? In 4.3 he says explicitly that no strategy is better than another according to our current understanding of epistemics and decision theory, that the EA project needs rethinking of its epistemic and decision-theoretic foundations, and that he'd love to be wrong. So he isn't claiming the door is closed.
But look at which door he leaves open. The provisionality is about our epistemics and decision theory possibly improving. It isn't about evidence accumulating. And those come apart. If the load-bearing imprecision is the evaluative sort, then a better decision theory might change the verdict, while no amount of further empirical work would.
Which is the closest thing to a practical upshot I have, and it's why I'm submitting this as a constructive response rather than a critique. 4.3 says the EA project needs rethinking of its epistemic and decision-theoretic foundations, and leaves it there. The reading above narrows that a lot. If the imprecision doing the blocking is evaluative rather than mechanistic, then the work that could move things is decision-theoretic and metanormative, not empirical. More careful modelling of AI development won't do it. Neither will better forecasting, however much it improves. That's not a comfortable answer and it isn't the one I expected to arrive at, but it tells someone deciding where to put their effort something more specific than "we need better foundations." It also suggests a question worth asking of any particular cause before asking anything else: not how wide the imprecision is, but what kind it is, because that determines whether anyone could ever close it.
~
A strong reply I can see to this is that the formalism was never meant to track origin, and it isn't the formalism's job to support claims about tractability. That's pretty fair. But 2.2 is doing a lot of work in the sequence, it's what licenses the everyday cases where unawareness doesn't bite, and it's origin-sensitive by construction. A formalism that can't express the distinction its own tractability criteria depend on seems worth noticing, even if the argument survives.
I'm also not certain the hypothesis-level and fine-grained split lines up as cleanly with the mechanistic and evaluative split as I've suggested. Some fine-grained imprecision is surely mechanistic, and some hypothesis-level imprecision surely isn't. The alignment is a tendency, not an identity, and I'd want to see it examined more carefully than I've managed here.