This competition entry has been selected for publication by the Forum team.
tl;dr
I consider the problem of “unawareness”: that “many possible consequences of our actions haven’t even occurred to us in much detail, if at all”, calling into question our ability to choose actions based on their possible consequences. I suggest modeling unawareness as arising from a lack of computational resources, specifically that even if I know relevant facts, the full implications of those facts may not be worked into my estimates because I haven’t had enough time to consider them.
I demonstrate that under this model unawareness can be bounded and estimated such that we can deliberate until the probability of unawareness reaches an arbitrarily low threshold. This leaves us with “mere” uncertainty within a model we understand ourselves to hold, and precise Bayesianism is a more natural fit there than it is for an agent who cannot yet state the possibilities she is uncertain over.
I further point out that this form of unawareness is well-studied in artificial intelligence, and that chess engines implement a form of the deliberation I propose here, lending empirical evidence to the view.

Our best estimate of v, the amount by which one action is better than another. As our estimate of v changes over time we can derive a prediction about how much our estimate is likely to change in the future. Note that this estimated uncertainty is not monotonic: if we receive an unexpectedly large change to our estimate (as we did at step six in this example) then the amount of uncertainty we have increases over time rather than decreasing. Nonetheless, if the rate of discovering sign-flipping considerations becomes low enough, we can become arbitrarily confident that one action is better than another (i.e. that v > 0), at least as estimated by our own (possibly flawed) model.

Unawareness in chess: a naive chess player may think that Bxd4 is the best move, as it wins a pawn. But a player who understands the “crucial consideration” that the knight is pinned to f3 realizes that this move is a blunder. We encounter DiGiovanni’s “pessimistic induction” - having seen one instance of a move flipping from good to bad, are we now likely to find even more examples? The answer is “yes” and we can in fact quantitatively estimate our level of unawareness. Chess engines use this estimate to determine when they should stop evaluating moves and instead go with their best guess.
Problem Statement and Summary of Resolution
From the perspective of traditional decision theory, certain types of unawareness should not exist. Suppose I know that people who stop eating beef will partially switch to eating smaller animals, and that smaller animals have more days of bad factory farming existence per calorie provided. If you tell me that this therefore means that encouraging people to stop eating beef will increase the number of factory farming days experienced by animals and therefore anti-beef campaigns are bad, I have learned nothing: no new evidence has arrived, and a Bayesian agent's credences are already closed under implication. Conditioning has no slot for this kind of update. Yet empirically, people do revise their estimates in response to purely theoretical argumentation of this kind, and it seems they are right to.
DiGiovanni diagnoses this as a problem of coarse-grained awareness: the outcomes I can conceive of are too coarse to support a well-defined expected value, and encountering the wild-animal argument reveals not just a missing consideration but evidence of how much my model omits. On his view no refinement rescues the comparison — each newly discovered consideration is an instance of a pattern (the pessimistic induction) suggesting the considerations still missing would swing the verdict again, so expected values remain severely imprecise and judgment should be suspended.
I suggest that this is better modeled as a computational constraint. The relevant facts were in my possession all along; what I lacked was the computation that combines them. My estimates are not the conditional expectations of an ideal reasoner but the running outputs of a bounded one, and being handed an argument is being handed the result of a computation I had not performed. Unawareness, on this view, is not a defect of the state space but ordinary logical non-omniscience.
This reframing suggests why unawareness may be more tractable than empirical cluelessness. The world is under no obligation to send us relevant evidence, but the order in which we generate considerations is not similarly indifferent: deliberation plausibly surfaces considerations in roughly size-biased order — the most important ones first, at least as measured by our own (possibly mistaken) lights. If consideration sizes arrive in size-biased order, the anticipated-instability term is dominated by a computable tail, and we can certify a stopping point: a stage of reflection past which the probability of a sign flip from any not-yet-considered consideration falls below a chosen tolerance. Unawareness then stops being an unbounded regress and becomes a budgeted term in the analysis.
What remains after this term is controlled is "mere" empirical cluelessness. Our uncertainty about the world may still be vast and our expected-value estimates correspondingly unstable. But once unawareness is reinterpreted as bounded computation, the case for exotic doxastic machinery weakens: the residual is uncertainty within a model we understand ourselves to hold, and precise Bayesianism is a far more natural fit there than it is for an agent who cannot yet state the possibilities she is uncertain over.
DiGiovanni states the following normative premise: To justify preferring action A over B on impartial altruistic grounds, we need to "expect" that our epistemically idealized self — the version of us that could aggregate all of A's and B's possible consequences — would assign A higher expected value. Call that idealized verdict V*. Let v be my current estimate of V* and v₁, v₂, v₃, … be a sequence of estimates produced by my ongoing deliberation.
Assume that a rational bounded agent cannot predict the direction of its own future updates-from-reflection, i.e. the sequence of estimates v₁, v₂, v₃,… is a martingale.
Define the anticipated instability R as the expected total (squared) future movement of my estimate under continued reflection — how much, by my own lights, my number is still going to swing. By the martingale property, R is also the conditional variance of V* around my current estimate v: my forecast of my own future mind-changes is my uncertainty about where reflection ends up.
Then the probability of a sign flip — that I currently rank A over B, yet after unbounded reflection would rank B over A — satisfies, with no distributional assumptions (this is Cantelli's inequality applied to the martingale limit):
P(flip) ≤ 1 / [1 + (|v| / √R)²]
We clearly know v, and so we have reduced the problem of estimating the probability of a sign flip to that of estimating R (the expected future movement of my estimate under continued reflection).
DiGiovanni notes:
past discoveries of insights that flipped the apparent sign of interventions are evidence that other such insights may exist
Taking this seriously allows us to estimate R, and therefore P(flip).
One way to model this is with a “stick-breaking” distribution. This allows us to estimate the remaining “tail” of updates from what we’ve observed so far. I won’t elaborate here (an LLM can explain) but it looks like this:

Our best estimate of v, the amount by which one action is better than another. As our estimate of v changes over time we can derive a prediction about how much our estimate is likely to change in the future. Note that this estimated uncertainty is not monotonic: if we receive an unexpectedly large change to our estimate (as we did at step six in this example) then the amount of uncertainty we have increases over time rather than decreasing. Nonetheless, if the rate of discovering sign-flipping considerations becomes low enough, we can become arbitrarily confident that one action is better than another (i.e. that v > 0).
Update Ordering
The key controversial claim here is that, even though we theoretically could get updates in any order, in practice we are likely to receive the largest updates earliest in our deliberation. Specifically: I claim that I can somewhat accurately predict which considerations are most likely to change my mind, even if I can't predict which considerations a more fully informed agent would tell me to prioritize.
Suppose I am deliberating about which of two AI safety research projects to pursue. My earliest considerations will be about things like how rapidly the project could complete or how useful the results will be for some theory of change. Only after very extended deliberation do I get to things like “which project seems cooler to my friend Bob.”
Now it may be the case that actually “which project seems cooler to my friend Bob” was the most important consideration (e.g. because 5 years from now Bob will be some influential politician), and in that sense my deliberation was incorrectly prioritized. But this is prioritization from outside my model - inside my (inaccurate) model of the world, I was correct to deprioritize this consideration.
I therefore claim that one can expect to somewhat rapidly converge to a position where one expects that further deliberation will not update their estimates about which of two actions is better. This leaves us with “merely” the uncertainty that comes from lacking empirical information about the world. And while that uncertainty may be vast, it is of a type that seems more appropriate for traditional Bayesian reasoning techniques.
Can we actually predict the considerations which we consider to be the most important?
If, as I claim, it is relatively easy to identify the considerations which will update our (possibly inaccurate) models of the world, then it should be rare to find instances where people had deeply considered an issue and only after deep deliberation realized a crucial consideration that they theoretically could have thought of earlier (because they had all the relevant facts and just hadn’t put them together).
It is hard to construct an entirely persuasive data set here, but I have written a critique of some examples people have used to claim the opposite. To the extent that you believe that it’s hard to find examples of deliberative altruists misestimating what the important considerations are (by their own lights), then you might consider the “stick-breaking” assumption of the model I propose here to be supported.
My anecdotal experience is that the most compelling examples of sign flipping considerations are ones where important empirical facts weren’t known at the time of decision, rather than ones where there was a computational constraint.
Example: Unawareness in Chess

Unawareness in chess: a naive chess player may think that Bxd4 is the best move, as it wins a pawn. But a player who understands the “crucial consideration” that the knight is pinned to f3 realizes that this move is a blunder. We encounter DiGiovanni’s “pessimistic induction” - having seen one instance of a move flipping from good to bad, are we now likely to find even more examples? The answer is “yes” and we can in fact quantitatively estimate our level of unawareness. Chess engines use this estimate to determine when they should stop evaluating moves and instead go with their best guess.
Chess is a convenient example for cluelessness as all “unawareness” is computational, in the sense that infinite computation could inform us of the complete consequences of any move. Given this, how do chess engines decide when to stop reasoning and go with their best guess?
The leading chess engine Stockfish uses several heuristics, notably including how frequently its guess of the best move has changed in its search process. This is essentially DiGiovanni’s “pessimistic induction”: the more frequently we encounter arguments which change the optimal strategy, the less willing we should be to go with our current best guess.

“Pessimistic induction” in chess: the engine deliberates for 16 iterations, and on the 16th extends deliberation because the best move switched from Nf3 to Bc4. It switches again on the 18th iteration back to Nf3. Then on the 20th iteration the best move has remained Nf3 for three iterations, so deliberation terminates and the engine chooses move Nf3.
These heuristics have formal justification which apply to any sort of probabilistic reasoning (including, as I argue here, unawareness), though they do benefit from a choice of parameters which was fine-tuned based on knowledge of chess (and we presumably lack similar knowledge about how to do good).
Limitations
This post deals with the problem of choosing an action given the information that you have available to you. It makes no claim that the information available to you will be good. It is entirely possible that your best guess about what will increase the value of the long-run future is quite bad.
Secondarily, even if we can bound the amount of unawareness we have given a certain amount of information, there is no guarantee that this bound will be particularly small. If every day we wake up and discover that there's some new consideration we haven't thought of which wildly updates our estimates of what to do, then there is no bound as to how confused about the world we can predict ourselves to be.
Conclusion
If we reframe unawareness as being about computational constraints instead of the “coarseness” of the hypothesis space, we can bound the probability of discovering a novel sign-flipping consideration. This doesn’t give us a guarantee that our choices will be good ex-post, but it does give a motivation for believing that standard bayesian decision theory tools like EV maximization are appropriate.
I don't think chess is a particularly good example. Given some finite amount of search, chess engines use heuristics to estimate the winning probability of some position (or otherwise give it a score), but as Winning isn't enough says:
Which is the case for chess, in which heuristics work as well as they do because they have been trained (in modern engines) on a large amount of games, so that whatever new situation they encounter can be considered to be in distribution. If this isn't the case, engines can fail spectacularly. And in reality we can't assume the future will be in distribution.
Thanks! Not sure I agree with "heuristic": Stockfish's stopping rule is a heuristic, but it's backed by a theorem (Cantelli) that holds with no distributional assumptions. What is distribution-dependent is the estimate of that feeds into it, and I agree that some skepticism is warranted there.
The chess analogy is meant to show the framework is implementable, not merely theoretical. That moves the question from "are expected values well-defined?" to "what is and how large is relative to it?" I consider that progress because the answer isn't uniformly "suspend judgment": some parameter values license acting, others clearly don't.