This competition entry has been selected for publication by the Forum team.
The unawareness sequence by Anthony DiGiovanni argues that impartial altruists face a deep challenge when trying to compare the expected value of different strategies. We are not merely uncertain about which outcomes will occur, but many relevant outcomes are too coarsely represented, or not represented at all, in our current thinking. DiGiovanni argues that, under such unawareness, precise expected-value comparisons are unjustified, and the imprecise comparisons we are left with are indeterminate. This in turn is taken to imply that we have no reason to choose one strategy over another from the standpoint of impartial altruism.
The sequence is a welcome call to reflection. It invites us to think more deeply about whether our formal representations of our beliefs are justified, and especially whether we have properly accounted for the problem of unawareness: hypotheses that we are currently not even aware of.
There is much that I agree with in the sequence. The extent of our uncertainty and unawareness is too rarely appreciated, and longtermist strategies and interventions tend to be endorsed with excessive confidence. The distinction between implementation- and outcome-robustness is important and further complicates our assessments, while the pessimistic induction from the history of sign-flipping considerations should give us serious pause about the sign of virtually any intervention.[1]
This essay will focus on the parts I disagree with or think need further development. Specifically, I will critique the inference from “we have large and imprecise uncertainty and unawareness” to “we have no reason to favor one strategy over another from the standpoint of impartial altruism.” This inference rests critically on the maximality rule, which roughly says that one option is favored over another only when it wins across all admissible ways of making our imprecise beliefs precise; otherwise, the comparison is simply indeterminate.
I will argue that this is implausible as an exhaustive criterion of justification (or preference or betterness), even if a version of the maximality rule might be a plausible account of full justification. In its place, I will defend a graded account on which justification comes in degrees, so that one strategy can be more justified than another even when neither is maximality-preferred. Building on this theoretical critique, I will end by arguing for a wager on the most justified strategies.[2]
Before discussing the maximality rule, it is useful to raise a foundational question about reasons-based choice that will serve as background for this discussion.
There is substantial disagreement about what it is to have good reasons for action, and there is much literature exploring that question.[3] My focus here will be narrower: whether the structure of reasons-based choice is best understood as categorical or graded.
The following illustrates what I mean: in a choice between A and B, should our judgment of whether A is better or more justified than B be understood as categorical or graded, as the reasons for and against each option vary? For example, should our judgment be forced to fall into categories of either strictly better, exactly equal, or strictly worse, which may be represented with the values +1, 0, or −1; or should it be allowed to assume a broader range of graded judgments that fall between those strict categories?
This structural question applies whether we are concerned with justification, betterness, preferences, choiceworthiness, or the like: in each case, the vertical axis can be understood as representing our comparative assessment of A and B in terms of the relevant notion.
Importantly, the question is not about precision versus imprecision: the categorical model and the graded model can each be made more or less precise, such as by adding “fuzziness bars” of varying width around their central values. The question is instead whether changes in our reasons for or against a given option should ideally imply categorical or gradual changes in our assessments of comparative justification, betterness, and the like.
The reason I raise this question is that the maximality rule I will be critiquing is a categorical one: in pairwise comparisons between options, it implies that one option is strictly better than the other, the two options are exactly equal, or their relation is simply indeterminate, with no degrees in between.
This categorical feature is in no way unique to the maximality rule. For example, the main decision criteria for imprecise credences considered by Mogensen (2021) are all categorical in this broad sense: they determine discrete preference relations or choice statuses among options rather than representing degrees of comparative preference or justification.[4]
However, while such discrete or categorical preference relations (or betterness relations, decision criteria, etc.) are common and have the desirable feature of being simple, there are reasons to think that the categorical approach is not always the most plausible one. At the very least, it is worth exploring alternatives before settling on which approach is best.
To motivate this discussion, the following are some general reasons we might favor a graded approach, at least in some evaluative domains. First, we may find it plausible that adding a weak reason for a given option should at most lead to a gradual increase in our level of preference or justification for it. In particular, categorical jumps in our level of preference or justification in response to adding a weak reason seem arbitrary unless we can explain why the weak reason should carry such outsized significance.[5]
Second, if we use degrees to represent how our beliefs change in response to evidence, it seems natural to do the same for our evaluative stances in response to reasons, as reasons plausibly also count as incremental evidence for evaluative stances.[6] Do we have a principled reason to allow our beliefs to come in degrees while not allowing the same for our evaluative stances? If not, this asymmetry looks unmotivated.[7]
Third, we may favor a graded approach because it allows for a much wider range of possible judgments and preferences, as opposed to being restricted to a few rigid categories.[8] In this way, the graded approach has much in common with imprecise credences: it avoids an artificial narrowing of allowed positions, and it may track our underlying reasons better than a few sharp categories can.
Fourth, we may favor a graded approach because it can reflect action-guiding differences in evaluations and reasons-based choice that a categorical approach would erase. For instance, where a categorical approach might say that all comparisons are indeterminate, a graded approach may enable us to discriminate options based on their varying degrees of justification, and to then choose the option that seems most justified.[9]
The framework of imprecise probabilities adopted in the unawareness sequence involves representing our beliefs with a set, P, of plausible probability distributions, p ∈ P. This set is called a representor. By taking the expected value of a given action, A, for each p ∈ P, we in effect get a set of expected values of A.
The maximality rule says that action A is strictly preferred to action B if and only if the expected value of A is higher than the expected value of B for every p ∈ P.[10] If neither action has a higher expected value for every p ∈ P, and if their expected values differ for at least one p ∈ P, then it is simply indeterminate which action should be preferred.[11]
This formal relation does not by itself imply that all comparative impartial reasons disappear whenever its demanding condition fails. That stronger conclusion requires a further exhaustiveness thesis: unless one act is maximality-preferred, there is no impartial reason to favor it. It is primarily this further thesis that I dispute.[12]
My criticism therefore centers not on maximality as a possible sufficient condition for strict preference, but on the categorical jump from the absence of maximality preference to total indeterminacy, with no weaker or graded reasons in between. This all-or-nothing framework is prone to blanket indeterminacy and non-discernment by design, which in turn leads to implausible implications, as I will argue below. I will first present some broad structural objections to the maximality rule, and then build on these with some practical objections.[13]
Suppose that all but one of the probability distributions in P favor A over B by an overwhelming margin, while a single outlier distribution favors B over A by the tiniest margin.[14] According to the maximality rule, this single outlier distribution is enough to make our preference between them completely indeterminate.[15]
Similarly, suppose that, in comparison to a hypothetical null action, action A has an imprecise expected value of [−ε, 1,000,000], where ε is arbitrarily small. The maximality rule would leave it completely indeterminate whether to favor A or the null action, and this holds true however large the upper endpoint may be.
Or, to phrase the case in more concrete terms, imagine that a group of sentient beings stands to be tortured for life, and you can either choose the null action that is guaranteed to prevent none of it, or choose action A, whose imprecise expected impact ranges from adding a nanosecond of torture to preventing all of it.[16] The maximality rule leaves the choice perfectly indeterminate and deems the null action rationally permissible.[17]
This veto power in favor of complete indeterminacy seems implausibly strong. Why should a single outlier distribution, or a single dissenting extreme point of P, have such decisive veto power in the face of an otherwise strong consensus in favor of A? Even if the dissent gives us reason not to fully prefer A, it still seems plausible to have some graded preference for A (for instance, a preference of degree 0.9), rather than to declare complete indeterminacy simply because that is all this categorical formalism allows.[18]
Building on the null-action comparison, say that you encounter an arbitrarily weak bit of evidence that leads you to update the imprecise expected value of A from [−ε, 1,000,000] to [ε, 1,000,000], where ε is arbitrarily small. This tiny update takes us from complete indeterminacy to a strict preference according to the maximality rule.[19]
Similar to the case above, it seems implausible for an arbitrarily weak bit of evidence to have such categorical significance, in effect pushing us over the sharp cliff of maximality. Why should this weak bit of evidence carry such categorical significance for our preferences or comparative judgments while other, far stronger pieces of evidence carry no significance whatsoever?
The maximality rule here commits us to a stark discontinuity: the strong pre-existing support for A has effectively zero weight in the comparative verdict before the update and decisive weight after it. Again, it seems more plausible to instead allow for a graded conception of preference or ex ante betterness that altogether avoids such categorical shifts in response to arbitrarily weak evidence.
The objections so far grant that the maximality rule at least delivers determinate verdicts. But this assumption itself is doubtful if the boundaries of our expected-value estimates are vague rather than sharp.[20]
Such vagueness creates a dilemma. Either we precisify the boundaries artificially, in which case the threshold between complete indeterminacy and strict preference rests on a stipulation that, as proponents themselves acknowledge, does not reflect our actual epistemic state. Or we let the boundaries remain vague, in which case it becomes indeterminate whether our preferences are indeterminate: a form of higher-order indeterminacy that deprives the maximality rule of the clear and simple verdicts that were among its main virtues.
A graded conception of preference or betterness can dissolve this dilemma and give more plausible answers in the face of vague credences: since betterness comes in degrees, and since these degrees can themselves be vague, vague endpoints in credences can be tracked by vague degrees of betterness. There is no sharp cliff whose location we must either stipulate or leave undefined.
For example, in a comparison between A and B, the picture might look roughly like this with a graded and vague conception of betterness:[21]
The arguments above show that the maximality rule has implausible implications in stylized cases. In this section, I will argue that the same is true in real-world cases. Specifically, the maximality rule faces a dilemma. Either the rule yields indeterminacy on every scale, in which case it is practically silent and seems implausible as a standard for outcome-based choice. Or it yields determinate verdicts on sufficiently small scales but becomes indeterminate as the scope or time horizon expands, in which case it implies an implausible discontinuity: a marginal widening of our uncertainty eventually produces a categorical shift from strict betterness to complete indeterminacy.
The following two subsections develop these horns in turn. I do not take a position on how small a scale is required before maximality yields a determinate verdict, or whether it ever does. The objection is that each horn would speak against the rule, and that graded approaches seem to do better under each horn.
Why might one think that maximality implies universal indeterminacy, including on small scales where indeterminacy seems implausible? One set of reasons derives from “known unknowns”: hypotheses that we are aware of but highly uncertain about. For example, we do not know whether the simulation hypothesis is correct, and it could introduce substantial uncertainty even in small-scale outcomes.
Similarly, there is a vast space of “unknown unknowns”: the possibilities we are unaware of. This unawareness extends across many levels and also includes deep forms of ontological uncertainty that go beyond the framework of established physics. This uncertainty may also be relevant to small-scale outcomes. With such known unknowns and unknown unknowns, one might hold that we are unlikely to ever reach the perfect unanimity required by the maximality rule.[22]
Moreover, if we grant that our imprecise credences have vague endpoints, this only seems to strengthen the case for small-scale indeterminacy under the maximality rule. First, exotic known unknowns and unknown unknowns are arguably among the considerations that make the endpoints of our imprecise credences the most vague and hazy, since it is so difficult to estimate their significance.
Second, if we want to be firm in applying the maximality rule, we should presumably interpret the vague endpoints of our imprecise credences in an expansive sense, so as to avoid excluding admissible credence functions. Yet the more expansively we draw these boundaries, the more dissenting distributions we include, and the harder it becomes to reach the perfect unanimity required by the maximality rule.
A proponent of maximality might at this point decide to bite the bullet: yes, the maximality rule leaves our evaluative stance perfectly indeterminate and thus leaves us clueless even with respect to small-scale goals, but that is simply the predicament in which we find ourselves.[23] What would be the alternative?
But the point is that there is a broad family of alternatives: graded approaches to betterness. To see how these approaches may diverge from maximality even on a local scale, consider a choice between A and B, where A is a commonsensically sound strategy for avoiding serious injury, such as taking ordinary safety precautions, while B is a commonsensically inferior strategy, such as disregarding those precautions.
If we grant that we cannot reach perfect unanimity about which strategy is best for avoiding serious injury, due in part to exotic possibilities and unawareness, the picture looks roughly as shown below: although our reasons overall seem to lean considerably toward A, the maximality rule still implies total indeterminacy, whereas a graded approach would imply that A has a high degree of ex ante betterness.
The advantage of the graded approach is that, even if we cannot reach strict betterness in the strong sense of perfect unanimity across all admissible credence functions, we can still obtain meaningful degrees of betterness that reflect our available reasons and evidence. This seems like a more plausible approach, especially when our reasons overall point considerably more toward one option than another. In contrast, if the maximality rule implies total indeterminacy even for the simplest goals, this is arguably a reductio of maximality as an exhaustive criterion for outcome-based choice.
Let us turn to the second horn: maximality does not imply indeterminacy on small scales, but only as the scope or time horizon expands. Even if we grant this premise, the maximality rule still leads to implausible implications.
The problems we would face in this case are real-world versions of the structural problems we saw in the stylized cases above. Say that maximality does not imply indeterminacy about action A compared to a null action on very small scales, but as the time horizon increases and A’s range of expected values grows wider, indeterminacy eventually emerges.
This picture faces at least two major problems. First, we have the categorical jump from strict ex ante betterness to total indeterminacy. This jump seems implausible because a slight extension of the time horizon, and thus a slight widening of A’s EV range, suddenly flips our verdict from strict betterness to total indeterminacy: a categorical shift with no correspondingly significant change in our epistemic state or underlying reasons. This is the stark discontinuity from before, now running in reverse: the same body of support effectively carries total weight at one time horizon and none at all just beyond it.
Second, we face the problem of vague endpoints: when the endpoints of our imprecise credences are vague, where do we choose to make the categorical jump from strict betterness to total indeterminacy? This choice can seem arbitrary, which makes the categorical jump seem even less plausible: not only do we make a categorical shift without a correspondingly large change in our underlying reasons, but we even seem to make this abrupt shift at an arbitrary point.
Both problems arise from maximality’s categorical conception of betterness, and this example shows that, on the second horn, these problems also afflict the maximality rule in practice. Once again, a more plausible view is that the ex ante betterness of A gradually declines as the time horizon increases.[24]
In summary, the maximality rule is not plausible as an exhaustive criterion for comparative justification or betterness. Even if maximality is necessary for full betterness, it does not follow that it is necessary for one option to be better than another to some degree, or for us to have more reason to favor it. Treating maximality as exhaustive collapses potentially large differences in the direction and strength of our reasons into a single category of indeterminacy. A more plausible account would preserve those differences while allowing residual indeterminacy and imprecision.[25]
Graded accounts of justification, betterness, and related notions form a broad class. That is, just as there are many ways to formulate categorical rules for strict betterness, such as maximality and Γ-maximin, there are also many ways to formulate graded approaches.[26] We should thus avoid confusing any specific graded approach and its individual plausibility with the wider class of graded approaches and its plausibility as a whole.
My core claim here is that graded accounts as a class are more plausible than categorical accounts that jump abruptly between strict betterness and complete indeterminacy. More precisely, my claim is that there are accounts within the broad class of graded approaches that are more plausible than any of those abrupt accounts. The arguments above support this claim independently of any specific graded account, since they target the broad shape of the categorical approach: the abrupt jump itself.[27]
However, I do not claim to have identified the most plausible graded accounts; that would be a task for further research. The specific approaches I present below are merely simple examples of graded approaches that, in my view, already seem more plausible than the maximality rule.
For the two illustrative approaches below, assume that the relevant EV ranges are bounded.[28] Say that we compare A to a hypothetical null action whose expected value is 0 for each p ∈ P. This gives us a range of expected values relative to the null action: [inf(A), sup(A)]. A natural way to define the degree of betterness or comparative justification of A relative to the null action when the interval straddles 0 (i.e., inf(A) < 0 < sup(A)) is then:
D(A) := (inf(A) + sup(A)) / (sup(A) − inf(A))
Otherwise, when the interval lies above or below 0, we simply say that A is fully better or fully worse than the null action, or better to degree 1 or −1, depending on which side of 0 the interval falls on.[29]
The formula above is equivalent to:[30]
D(A) = (share of the EV range lying above 0) − (share lying below 0) = midpoint / half-width[31]
Consequently, when the EV range is mostly above 0, A has a positive degree of betterness compared to the null action, whereas it has a negative degree of betterness when it is mostly below 0.
For intervals that straddle 0, this approach has the advantage that D smoothly approaches 1 as the infimum approaches 0 from below, and approaches −1 as the supremum approaches 0 from above.[32] It thereby avoids any discontinuity at the boundary between graded and full betterness, the point at which the maximality rule instead jumps abruptly between total indeterminacy and strict preference.
Similarly, for such intervals, since D changes continuously as the endpoints of the EV range vary, endpoint vagueness is reflected in graded rather than categorical variation in the verdict: when the endpoint vagueness or variation is small relative to the interval’s width, the resulting variation in D is also small. There are thus no categorical jumps hanging on arbitrary precisifications.[33]
Another advantage is that the approach conservatively extends the maximality rule: whenever maximality yields strict betterness or strict inferiority relative to the null action, the graded approach agrees and assigns D = 1 or D = −1. Its distinctive contribution is only to provide finer-grained verdicts in cases that maximality leaves indeterminate. Instead of collapsing everything into the categories of strict betterness or total indeterminacy, this graded index restores credence-sensitive discriminations among options.
In particular, we gain an index that is sensitive to mild sweetenings in a natural cumulative fashion. Likewise, this approach can in principle favor beneficial near-term interventions even when the long-term effects are unclear. Suppose that intervention A adds the same direct benefit d > 0 relative to a null option for every p ∈ P, while its flowthrough EV range is symmetric at [−h, h]. The total EV range of A is then [d − h, d + h], and when this range straddles 0, A’s degree of betterness is D(A) = d / h.[34] Widening the flowthrough uncertainty thus gradually weakens the case for the intervention without erasing it altogether.[35]
It is worth reiterating that D need not be taken to yield a fully precise value in every case. Just as we can admit vagueness in the endpoints of our imprecise credences, we can admit corresponding vagueness in the value yielded by D. This makes it difficult to say whether a given option is better than the null action when D is close to 0, which seems quite plausible: cases where D ≈ 0 are exactly the kinds of cases where betterness or comparative justification may be too unclear or too vague to discriminate either way. Yet the virtue of this approach is that we can grant limited discrimination in those cases without granting total indeterminacy across the entire range.[36]
The approach mentioned above is too simple to guide us in general: it is only designed to assess betterness for one action compared to a null action, for which it seems workable. Yet we cannot reasonably rank all actions based on this betterness index alone, since such an approach would have implausible implications similar to those raised against maximality earlier. For example, it would imply that an action with EV range [ε, 2ε] would be better than an action with EV range [−ε, 10100]. To rank multiple options simultaneously, we want an approach that is globally sensitive to the expected values of our options.
Thus, to construct a more plausible general approach, we could opt for the following, inspired by Clifton (2025b). For an action A, let U(A) be the interval spanned by the expected values of A, and let M(U(A)) be the midpoint of U(A). For two actions A and B:
where h and h are the half-widths of U(A) and U(B), respectively.[37]
The midpoint-based sign of CD determines the direction of the graded evaluation, while its absolute magnitude represents the strength. When comparing any two options, this strength may help inform how strongly we wager on the favored option and how much priority we assign to gathering further information.
CD is a generalization of the simple index constructed above, and it therefore shares many of its desirable features. Moreover, this approach has the advantage of mitigating a problem raised in Clifton (2025b), where an arbitrarily small sweetening produces an abrupt reversal in strict preference. Under the present approach, for any fixed pair of options, an arbitrarily small sweetening cannot take us from a full preference for one option to a full preference for the other; it can at most take us from a full preference for one to a very slight preference for the other.[38]
We might object that the approach outlined above is arbitrary. Specifically, with a range of expected values, why privilege the midpoint? Why reintroduce what appears to be arbitrary precision?
First, we should be clear that graded accounts in general need not involve midpoint ordering, and the case for graded approaches advanced here does not depend on it. There may well be more plausible graded accounts.
As for the plausibility of midpoint ordering, we may argue that it finds some weak support in a symmetry intuition.[39] More elaborately, we might argue that the midpoint is the least biased toward either of the endpoints: ordering options by a point above or below the midpoint would introduce a structural bias toward optimism or pessimism, respectively. Similarly, the midpoint can be defended as a form of worst-case error minimization, in that it is the point that minimizes the maximum absolute distance from any admissible expected value. Picking any other point within the interval may be said to arbitrarily increase the worst-case error on one side of the interval. Furthermore, midpoint ordering has the advantage that it preserves transitivity across comparisons.
These considerations do not establish midpoint ordering as uniquely plausible or correct, but they do lend it some tentative support. Moreover, midpoint ordering need not reintroduce exact precision, since vagueness in the endpoints can induce corresponding vagueness in the midpoint.
The points above may go some way toward addressing the arbitrariness objection against midpoint ordering. However, I believe a stronger reply to a proponent of maximality centers on the comparative arbitrariness of the maximality rule itself. Arbitrariness, and conversely defensible reasons and justifications, come in degrees. Thus, even if midpoint ordering entails some degree of arbitrariness, and even if better approaches exist, the approach outlined above may still overall be less arbitrary and more justified than the maximality rule.[40]
I have sought to make essentially this case throughout this essay with respect to maximality versus graded approaches in general: maximality’s categorical structure seems the more arbitrary and less plausible one overall. We can make the same case for the specific graded approach outlined above by looking at concrete examples.
For instance, in a choice between A and B where neither dominates the other and U(A) = [−1, 100] and U(B) = [−100, 1], maximality treats the comparison as perfectly indeterminate, whereas CD(A, B) ≈ 0.98.
Likewise, if a small update shifts U(A) from [0.1, 100] to [−0.05, 100] while U(B) is [0, 0.05], and neither dominates the other in the latter case, the verdict under maximality would go from a strict preference for A to complete indeterminacy. In contrast, the graded approach would go from a full preference to CD(A, B) ≈ 0.998.
Most starkly, consider a variant of the earlier torture case: we can choose either a null action, N, which is guaranteed to prevent no torture, or action A, whose expected impact ranges from adding a nanosecond of torture to preventing an arbitrarily large finite duration of it.[41] Maximality would yield complete indeterminacy while CD(A, N) = D(A) ≈ 1.
Which approach seems more arbitrary or implausible in light of these examples? While midpoint ordering can legitimately be accused of some measure of arbitrariness, maximality’s insistence on either strict preference or total indeterminacy, and the resulting abrupt jumps between these verdicts, seem more arbitrary overall.
Maximality privileges a poorly justified unanimity threshold, discards information about the magnitudes of the expected values, and can amplify minor arbitrariness in the specification of P into categorical differences in verdicts. More generally, it seems more arbitrary to refuse to make credence-sensitive verdicts whenever unanimity fails than to make graded verdicts that track our beliefs as closely as possible.
In the spirit of the objection above, we can ask: why privilege maximality and its categorical structure? Given the considerations outlined above, it seems that the burden is on the proponent of maximality to explain why it is superior to graded approaches that avoid many of maximality’s implausible features.
The framework presented in the unawareness sequence goes much of the way toward the broadly graded approach defended here: it embraces degrees and non-sharpness at the level of empirical beliefs. Yet its decision rule lies at the opposite end of the spectrum: maximality admits of no degrees, which is precisely why it is vulnerable to abrupt jumps based on small changes or misspecifications. In this way, the decision rule does not harmonize smoothly with the non-sharpness inherent in the belief apparatus.
The approach suggested here is simply to go all the way with the non-sharpness already embraced in the treatment of empirical beliefs: just as beliefs do not have to be sharp, neither do comparative evaluations of betterness or justification. Our evaluations can then track the non-sharpness of our beliefs rather than collapsing it, as illustrated loosely below.[42]
Having made a case for broadly graded approaches, I will proceed to explore how we might apply such approaches to identify plausible practical recommendations despite deep uncertainty.
This exploration will necessarily be limited: identifying the most justified or plausible recommendations is a major project, and one cannot expect to make much progress on it within a single section. But hopefully what follows will still show that this is a practically fruitful direction and give some sense of how graded approaches can help us avoid cluelessness by providing impartial action guidance that is justified to some degree.
On the view defended here, we have reason to wager on the strategies that are most justified by our available evidence, broadly construed, while keeping those wagers proportionate to the strength of their justification. This need not mean forcing a precise ranking of all options or betting everything on a single best guess. Our wagers can remain diversified, responsive to new evidence, and aimed largely at improving our estimates.
The main objection to wagering presented in the sequence is that any evidence or proposal we might wager on is too weak to overcome indeterminacy among our options. Our outlook is insensitive to mild sweetening, and hence we are stuck in total indeterminacy, which warrants no wagers.[43]
This objection assumes maximality. More generally, it assumes the kind of categorical decision rule that collapses a broad set of evidential states into complete and undifferentiated indeterminacy. If we grant such a decision rule, and further grant that our epistemic situation places us firmly within the rule’s zone of total indeterminacy, this objection succeeds. Yet under graded approaches, this wager-blocking conclusion does not follow.
I will focus on temporally extended strategies rather than one-time actions, since we have a greater chance of discriminating between the expected impacts of strategies than between those of individual actions. A single action tends to make less of a difference than many actions that all push in the same broad direction and that may compound over time. Thus, even if we cannot distinguish between the expected impacts of any two actions, we may still be able to distinguish the expected impacts of at least some strategies.
The illustration above shows the respective impacts of two strategies separating completely. However, what we need to make graded evaluations is merely that they diverge to some discernible degree, which is a much weaker condition.
Similarly, practical guidance requires only that we can distinguish the expected value of some strategies, not necessarily all. Even if most strategies are mutually indistinguishable in terms of their expected value, there may still be some that we can distinguish as better or worse than others in expectation.
Taken together, all we need for practical guidance is that some strategies seem better than others to some discernible degree.
What are the most plausible candidates for meeting these relatively modest conditions?[44] While it is difficult to say which candidates are most plausible or have the greatest comparative justification, the following broad strategies are each plausible contenders for having a relatively high degree of justification and thus for being worth prioritizing.[45]
Improve our comparative estimates
A promising strategy is to pursue targeted research to improve our comparative estimates and decision-making. This fits naturally with the graded approach: rather than relying on maximality and being stuck in complete indeterminacy, we can start out with tentative estimates of comparative justification and gradually refine these estimates.
This refinement can occur at many levels. For example, we can improve our views on:
Research that helps improve our views at any of these levels can plausibly help improve our graded estimates of which decisions are better justified, thereby addressing a key practical bottleneck.[46]
Our room for progress seems large across many, if not all, of these levels, and there are reasons to think that we can make substantial progress with further research. In particular, for many of these questions, the lines of inquiry addressing them are quite young and neglected, and past work has made significant progress on at least some of them. Moreover, it seems plausible that future AI systems can help us make meaningful progress at each of these levels and potentially uncover many crucial considerations. This further supports the case for wagering on such research.
Importantly, progress here is not merely a matter of honing our pre-theoretic intuitions. Our comparative estimates can also incorporate empirical evidence, formal models, and elaborate considerations that can be explicitly stated and refined. Unlike vague intuitions, such structured estimates might become vastly better as we accumulate stronger evidence and insights over time.[47]
Work on these questions may also be unusually time-sensitive: insights and refined estimates gained earlier can inform a greater number of subsequent decisions, including decisions that may be irreversible.
Expand flexible resources
A complementary strategy is to grow and develop flexible resources that can contribute to altruistic work. In concrete terms, this prominently includes money and people who are willing and able to contribute effectively to altruistic aims. By extension, it also includes upstream resources that support funding growth and help attract and develop more effective contributors.
This strategy is supported by various considerations. Flexible resources such as funding and competent people are core ingredients for impact: they can help achieve a broad range of goals, including the above-mentioned goals of improving our comparative estimates and decision-making. A further advantage is their option value: they can later be deployed toward goals that our better-informed future selves or altruistic successors deem worth pursuing.[48]
More generally, both the strategies of improving our estimates and expanding flexible resources have several shared advantages. They both seem valuable across a wide range of plausible futures and require no commitment to a narrow causal story about which specific interventions will prove best. Furthermore, the two strategies can be mutually reinforcing, as gains in each tend to raise the returns to the other: better estimates can guide the use of flexible resources, while expanded resources can support further improvements in those estimates.
These strategies also find some comparative support in the pessimistic induction. That is, the pessimistic induction from our track record of discovering sign-flipping considerations does have dampening force across the board, but this force does not apply evenly. It weighs most strongly against strategies whose value depends on a particular causal story surviving the next sign-flipping discovery, and least strongly against flexible strategies whose value largely consists in making such discoveries and being better positioned when they occur.
Likewise, if we look at our actual history of encountering sign-flipping considerations, we seem to find more considerations that flip the sign of specific large-scale interventions, and comparatively fewer considerations that flip the sign of low-footprint capacity-building such as improving our estimates and expanding flexible resources. Indeed, insofar as this historical record suggests that further important considerations can readily be discovered, it arguably provides some positive graded support for improving our comparative estimates and building flexible capacities that can be deployed in light of better information.[49]
One might worry that the two broad strategies outlined above are too general to provide concrete guidance. Yet each can be broken down into more specific priorities. Improving our estimates points toward concrete research questions that can be ranked, and iteratively re-ranked, by their apparent tractability and importance. Similarly, expanding flexible resources can draw on established insights about growing funding, attracting contributors, and developing talent.
Pursue best-justified interventions
If our research on comparative estimates identifies promising interventions, we can wager on these in proportion to their degrees of justification. In practice, this may imply a portfolio that combines comparatively robust interventions to reduce near-term suffering with more speculative but potentially high-impact efforts to improve the long-term future.
It is unclear what these interventions will turn out to be. From our current vantage point, some plausible candidates include efforts to make future AI systems robustly cooperative and to strengthen norms and institutions for peaceful bargaining and conflict resolution among powerful actors with divergent values.[50] Other candidates include efforts to differentially develop the capacities and dispositions of advanced AI systems such that they can safely support the first two strategies: improving comparative estimates and building flexible resources for altruistic work.[51]
These examples are necessarily tentative. The broader point is that graded comparison can support proportionate wagers on promising interventions while leaving room to revise our estimates and wagers as new evidence emerges. Our confidence in any particular intervention may remain limited, yet the broader strategy of placing revisable wagers on the best-justified options may itself be comparatively well justified.[52]
The case for cluelessness presented in the unawareness sequence rests on the maximality rule. I have argued that this rule is implausible as an exhaustive criterion of comparative justification. When maximality fails, we can still have some degree of impartial reason to favor one option over another.
Rejecting maximality does not automatically imply that we avoid cluelessness, but it does make cluelessness more difficult to defend. In particular, it is much harder to show that we are clueless under graded alternatives to maximality, especially when we consider self-refining strategies focused on improving our views and expanding our resources. It also becomes much harder to make a case against wagering, since that case appears to rest largely on categorical indeterminacy.
More broadly, I have argued that the sequence has not identified, much less undermined, the most plausible approach to impartial decision-making, whatever that turns out to be. There is still much work to be done in finding the least arbitrary, best-justified approach.
That work is itself among the wagers available to us. The practical upshot of a graded view is that we have reason to wager on the strategies that seem best justified, including those aimed at improving our estimates.
Improving our comparative estimates and expanding flexible resources are both forms of relatively low-footprint capacity-building, which the sequence criticizes in a dedicated section.[53]
Low-footprint capacity-building has plausible benefits: it “could make our successors discover, or more effectively implement, interventions that are net-positive from their epistemic vantage point.” Yet it might also prevent future interventions that are more beneficial in expectation under some admissible precisifications of our present beliefs. Under maximality, this is enough to defeat a strict preference for low-footprint capacity-building.
However, as I have argued above, a failure to satisfy maximality does not imply that a strategy cannot still have a higher degree of justification or ex ante betterness than another. To defeat that more modest standard, it is not enough to argue that there are some realistic downsides or that some plausible distributions favor another strategy.
The objection cites Appendix C in the sequence, in which a key passage reads:
Suppose that if you currently do nothing, your future self might take some action at “crunch time” that is neither robustly positive nor negative (from their epistemic perspective). Whereas, if you instead think more about how to have a robustly positive impact (see “low-footprint Capacity-Building” in the final post), your future self will take an action that’s robustly positive — but not robustly better than the default action. Then with respect to UEV [combined with maximality], you don’t have a reason to think more rather than do nothing.[54]
The illustration below shows what this scenario could in principle look like, where some plausible distribution favors the resulting default action over the robustly positive action:
The claim that “with respect to UEV, you don’t have a reason to think more rather than do nothing” need not follow under a graded approach. Moreover, even maximality does not itself say “you don’t have reason.” As a formal rule, maximality merely says that there is no strict preference between the two options and that both are rationally permissible.[55] It does not make explicit claims about fully capturing our underlying reasons.[56] One might interpret it that way if one takes the rule to exhaust our impartial reasons. But whether it does is precisely what is at issue. I have tried to argue that the maximality rule does not fully capture our underlying reasons.[57]
The sequence also cautions against “thinking of our future selves as perfectly coherent extensions of our current selves.” This is a fair caution, as our future selves may change in important ways, which is a relevant concern for the strategies proposed above. However, the case for low-footprint capacity-building need not assume that our future selves will be perfectly coherent extensions of our current selves; a reasonably high degree of continuity may suffice.
Furthermore, acknowledging that our future selves may change in significant ways does not imply deep uncertainty about maintaining a reasonably high degree of continuity in our altruistic commitment. How much such continuity we should expect is a complicated question, but it is not in the domain of far-future uncertainty. Rather, it is a question of psychology about which we have decent evidence.
For instance, longitudinal research finds substantial rank-order stability in core personal values across adulthood, including self-transcendence values such as benevolence and universalism, with self-transcendence values tending to increase with age.[58] Consistent with this, there are many examples of individuals who have sustained a high level of altruistic commitment for decades.[59]
More directly, and most relevant to our individual case, we may draw some confidence from the duration of our own altruistic commitment so far, which for some of us has been our entire adult lives or even longer. If that pattern has been highly consistent in the past, this gives us some reason to expect further consistency in the future.
Finally, it is worth noting that just as our altruistic commitments and values might change for the worse, they might also improve. Most of us are not currently at some global optimum of altruistic commitment and ideally reflected values from which we can only decline, or at best hold steady. Further improvement in these respects is likely possible, and the plausibility of such improvement may find support in our past experience of moral development, as well as some modest support in the psychological literature.[60] From this perspective, the seemingly unrealistic ideal of perfectly coherent extensions of our current selves might even be aiming too low.
The wagers described above focus on strategies at the practical level under graded approaches. Yet there are also many other levels at which wagers could be made, including various meta-levels.
The sequence brings up one example in its Appendix A: what it calls a meta-epistemic wager on precise subjective Bayesianism. That is, within a broader outcome-focused framework, residual weight on precise credences might warrant wagering on precise credences as our epistemic framework. The sequence considers this wager and raises some challenges for it.[61]
However, we can also make wagers within the imprecise framework: even if we place all our weight on imprecise credences, we can still wager on approaches that (sometimes) yield comparative verdicts over those that do not. In particular, we may wager on graded approaches over maximality, even if we give substantial weight to the latter. And note that maximality itself would not oppose such a wager: if the wager is not dominated by another strategy, it is perfectly permissible under maximality.
The illustration above depicts a situation in which we give roughly equal weight to maximality and some plausible graded approach, which I consider fairly generous to maximality in light of the objections raised earlier. But in any case, maximality is not the only option given imprecise credences: there are many alternative approaches, both categorical and graded ones, and it seems difficult to justify giving exclusive weight to maximality.
The wagers above are made within an outcome-focused moral framework. But we can also wager at a prior level, with the aim of making that framework workable. For example, if we find an outcome-focused moral framework to be the most plausible one from the standpoint of moral reflection, and if we are unsure how to make it work so that it provides justified verdicts, we may wager on a research program aimed at grounding and developing this framework. This might overall be the best-justified option at that prior level.
As these examples illustrate, there are multiple levels at which we can wager with the aim of securing justified practical verdicts under an outcome-focused framework. These wagers will not all yield exactly the same recommendations, which makes it worth clarifying which level we are mainly wagering at (if we favor one level in particular).
At the same time, the wagers outlined above may still yield a substantial level of convergence. In particular, they plausibly all broadly recommend foundational research and capacity-building work that enables such research.
To the extent that these practical recommendations find convergent support in several distinct wagers, this may give us further reason to pursue them, all things considered.[62]
Alvarez, Maria, and Jonathan Way. 2024. “Reasons for Action: Justification, Motivation, Explanation.” In The Stanford Encyclopedia of Philosophy, edited by Edward N. Zalta and Uri Nodelman. https://plato.stanford.edu/archives/fall2024/entries/reasons-just-vs-expl/.
Armon, Cheryl, and Theo L. Dawson. 1997. “Developmental Trajectories in Moral Reasoning across the Life Span.” Journal of Moral Education 26 (4): 433–53. https://www.tandfonline.com/doi/abs/10.1080/0305724970260404.
Armon, Cheryl, and Theo L. Dawson. 2003. “The Good Life: A Longitudinal Study of Adult Value Reasoning.” In Handbook of Adult Development, edited by Jack Demick and Carrie Andreoletti, 271–300. Springer. https://link.springer.com/chapter/10.1007/978-1-4615-0617-1_15.
Bostrom, Nick. 2003. “Are You Living in a Computer Simulation?” The Philosophical Quarterly 53 (211): 243–55. https://simulation-argument.com/simulation/.
Broome, John. 2004. “Reasons.” In Reason and Value: Themes from the Moral Philosophy of Joseph Raz, edited by R. Jay Wallace, Philip Pettit, Samuel Scheffler, and Michael Smith, 28–55. Clarendon Press. https://stafforini.com/works/broome-2004-reasons/.
Cailloux, Olivier, and Yves Meinard. 2020. “A Formal Framework for Deliberated Judgment.” Theory and Decision 88 (2): 269–95. https://arxiv.org/pdf/1801.05644.
Clare, Stephen. 2023. “Great Power Conflict.” Problem profile. 80,000 Hours. https://80000hours.org/problem-profiles/great-power-conflict/.
Clifton, Jesse. 2025a. “Reasons-Based Choice and Cluelessness.” Jesse’s Substack, February 7. https://jesseclifton.substack.com/p/reasons-based-choice-and-cluelessness.
Clifton, Jesse. 2025b. “Just Take the Midpoint?” Jesse’s Substack, June 23. https://jesseclifton.substack.com/p/just-take-the-midpoint.
Colby, Anne, Lawrence Kohlberg, John Gibbs, and Marcus Lieberman. 1983. “A Longitudinal Study of Moral Judgment.” Monographs of the Society for Research in Child Development 48 (1/2): 1–124. https://www.jstor.org/stable/1165935.
Cotton-Barratt, Owen, and Lukas Finnveden. 2026. “AI for AI for Epistemics.” Forethought. https://www.forethought.org/research/ai-for-ai-for-epistemics.
Dafoe, Allan, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R. McKee, Joel Z. Leibo, Kate Larson, and Thore Graepel. 2020. “Open Problems in Cooperative AI.” arXiv:2012.08630. https://arxiv.org/abs/2012.08630.
Dancy, Jonathan. 1993. Moral Reasons. Blackwell. https://philpapers.org/rec/DANMR.
Dietrich, Franz, and Christian List. 2013. “A Reason-Based Theory of Rational Choice.” Noûs 47 (1): 104–34. http://www.franzdietrich.net/Papers/DietrichList-ReasonBasedRationalChoice.pdf.
Dietrich, Franz, and Christian List. 2017. “What Matters and How It Matters: A Choice-Theoretic Representation of Moral Theories.” The Philosophical Review 126 (4): 421–79. https://philpapers.org/archive/DIEWMA.pdf.
DiGiovanni, Anthony. 2025a. “Should You Go with Your Best Guess? Against Precise Bayesianism and Related Views.” EA Forum. https://forum.effectivealtruism.org/posts/NKx8sHcAyCiKT723b/should-you-go-with-your-best-guess-against-precise.
DiGiovanni, Anthony. 2025b. “2. Why Intuitive Comparisons of Large-Scale Impact Are Unjustified.” EA Forum. https://forum.effectivealtruism.org/posts/qZS8cgvY5YrjQ3JiR/2-why-intuitive-comparisons-of-large-scale-impact-are.
DiGiovanni, Anthony. 2025c. “3. Why Impartial Altruists Should Suspend Judgment under Unawareness.” EA Forum. https://forum.effectivealtruism.org/posts/rec3E8JKa7iZPpXfD/3-why-impartial-altruists-should-suspend-judgment-under.
DiGiovanni, Anthony. 2025d. “4. Why Existing Approaches to Cause Prioritization Are Not Robust to Unawareness.” EA Forum. https://forum.effectivealtruism.org/posts/pjc7w2r3Je7jgipYY/4-why-existing-approaches-to-cause-prioritization-are-not-1.
Eisenberg, Nancy, Ivanna K. Guthrie, Amanda Cumberland, Bridget C. Murphy, Stephanie A. Shepard, Qing Zhou, and Gustavo Carlo. 2002. “Prosocial Development in Early Adulthood: A Longitudinal Study.” Journal of Personality and Social Psychology 82 (6): 993–1006. https://pubmed.ncbi.nlm.nih.gov/12051585/.
Evans, Geoffrey, and Anja Neundorf. 2020. “Core Political Values and the Long-Term Shaping of Partisanship.” British Journal of Political Science 50 (4): 1263–81. https://www.cambridge.org/core/journals/british-journal-of-political-science/article/abs/core-political-values-and-the-longterm-shaping-of-partisanship/D747688A17710DEDF73BC3CB69480056.
Fishburn, Peter C. 1970. “Utility Theory with Inexact Preferences and Degrees of Preference.” Synthese 21 (2): 204–21. https://link.springer.com/article/10.1007/BF00413546.
Hájek, Alan, and Wlodek Rabinowicz. 2022. “Degrees of Commensurability and the Repugnant Conclusion.” Noûs 56 (4): 897–919. Reprinted in The Philosophers’ Annual, vol. 41. https://pgrim.org/philosophersannual/41articles/hajekrabinowicz-degrees.pdf.
Hájek, Alan, and Michael Smithson. 2012. “Rationality and Indeterminate Probabilities.” Synthese 187 (1): 33–48. https://link.springer.com/article/10.1007/s11229-011-0033-3.
Herlitz, Anders. 2019. “Nondeterminacy, Two-Step Models, and Justified Choice.” Ethics 129 (2): 284–308. https://www.journals.uchicago.edu/doi/abs/10.1086/700032.
Ho, Lewis, Joslyn Barnhart, Robert Trager, Yoshua Bengio, Miles Brundage, Allison Carnegie, Rumman Chowdhury, et al. 2023. “International Institutions for Advanced AI.” arXiv:2307.04699. https://arxiv.org/abs/2307.04699.
Horty, John F. 2012. Reasons as Defaults. Oxford University Press. https://academic.oup.com/book/9900.
Kearns, Stephen, and Daniel Star. 2009. “Reasons as Evidence.” In Oxford Studies in Metaethics, vol. 4, edited by Russ Shafer-Landau, 215–42. Oxford University Press. https://philpapers.org/archive/KEARAE.pdf.
Knight, Carl. 2023. “Reflective Equilibrium.” In The Stanford Encyclopedia of Philosophy, edited by Edward N. Zalta and Uri Nodelman. https://plato.stanford.edu/archives/win2023/entries/reflective-equilibrium/.
Knutsson, Simon. 2021. “Many-Valued Logic and Sequence Arguments in Value Theory.” Synthese 199 (3–4): 10793–10825. https://link.springer.com/article/10.1007/s11229-021-03268-4.
Kollin, Sylvester, Jesse Clifton, Anthony DiGiovanni, and Nicolas Macé. 2025. “Bracketing Cluelessness.” Working paper, Center on Long-Term Risk, September 8. https://longtermrisk.org/files/Bracketing_Cluelessness.pdf.
Lewis, Gregory. 2021. “Complex Cluelessness as Credal Fragility.” EA Forum, February 8. https://forum.effectivealtruism.org/posts/Q3ZBt3X8aeLaWjbhK/complex-cluelessness-as-credal-fragility.
Li, Duo, Yuan Cao, Bryant P. H. Hui, and David H. K. Shum. 2024. “Are Older Adults More Prosocial Than Younger Adults? A Systematic Review and Meta-Analysis.” The Gerontologist 64 (9): gnae082. https://academic.oup.com/gerontologist/article/64/9/gnae082/7706145.
Lord, Errol, and Barry Maguire, eds. 2016. Weighing Reasons. Oxford University Press. https://academic.oup.com/book/1421.
Makins, Nicholas. 2023. “The Balance and Weight of Reasons.” Theoria 89 (5): 592–606. https://onlinelibrary.wiley.com/doi/10.1111/theo.12482.
Markovits, Julia. 2014. Moral Reason. Oxford University Press. https://academic.oup.com/book/9312.
Matsumoto, Yoshie, Toshio Yamagishi, Yang Li, and Toko Kiyonari. 2016. “Prosocial Behavior Increases with Age across Five Economic Games.” PLOS ONE 11 (7): e0158671. https://pmc.ncbi.nlm.nih.gov/articles/PMC4945042/.
Milfont, Taciano L., Petar Milojev, and Chris G. Sibley. 2016. “Values Stability and Change in Adulthood: A 3-Year Longitudinal Study of Rank-Order Stability and Mean-Level Differences.” Personality and Social Psychology Bulletin 42 (5): 572–88. https://journals.sagepub.com/doi/abs/10.1177/0146167216639245.
Mogensen, Andreas L. 2021. “Maximal Cluelessness.” The Philosophical Quarterly 71 (1): 141–62. https://academic.oup.com/pq/article-abstract/71/1/141/5828678. Working-paper version available at https://www.globalprioritiesinstitute.org/wp-content/uploads/Andreas-Mogensen_Maximal-cluelessness.pdf.
Mogensen, Andreas L., and David Thorstad. 2022. “Tough Enough? Robust Satisficing as a Decision Norm for Long-Term Policy Analysis.” Synthese 200 (1): 36. https://link.springer.com/content/pdf/10.1007/s11229-022-03566-5.pdf.
Montes, Ignacio, Enrique Miranda, and Susana Montes. 2014. “Decision Making with Imprecise Probabilities and Utilities by Means of Statistical Preference and Stochastic Dominance.” European Journal of Operational Research 234 (1): 209–20. https://bellman.ciencias.uniovi.es/~emiranda/stat-pref.pdf.
Montes, Ignacio, Enrique Miranda, and Susana Montes. 2017. “Imprecise Stochastic Orders and Fuzzy Rankings.” Fuzzy Optimization and Decision Making 16 (3): 297–327. https://bellman.ciencias.uniovi.es/~emiranda/FRV2.pdf.
Muehlhauser, Luke, and Anna Salamon. 2012. “Intelligence Explosion: Evidence and Import.” In Singularity Hypotheses: A Scientific and Philosophical Assessment, edited by Amnon H. Eden, James H. Moor, Johnny H. Søraker, and Eric Steinhart, 15–42. Springer. https://intelligence.org/files/IE-EI.pdf.
Nair, Shyam. 2021. “‘Adding Up’ Reasons: Lessons for Reductive and Nonreductive Approaches.” Ethics 132 (1): 38–88. Reprinted in The Philosophers’ Annual, vol. 41. https://philosophersannual.org/41articles/nair-adding.pdf.
Peterson, Johnathan C., Kevin B. Smith, and John R. Hibbing. 2020. “Do People Really Become More Conservative as They Age?” The Journal of Politics 82 (2): 600–611. https://digitalcommons.unl.edu/cgi/viewcontent.cgi?article=1118&context=poliscifacpub.
Pollerhoff, Lena, David F. Reindel, Philipp Kanske, Shu-Chen Li, and Andrea M. F. Reiter. 2024. “Age Differences in Prosociality across the Adult Lifespan: A Meta-Analysis.” Neuroscience & Biobehavioral Reviews 165: 105843. https://www.sciencedirect.com/science/article/pii/S0149763424003129.
Qureshi, Zershaaneh. 2025. “Using AI to Enhance Societal Decision Making.” Problem profile. 80,000 Hours. https://80000hours.org/problem-profiles/ai-enhanced-decision-making/.
Rabinowicz, Wlodek. 2012. “Value Relations Revisited.” Economics and Philosophy 28 (2): 133–64. https://www.cambridge.org/core/journals/economics-and-philosophy/article/abs/value-relations-revisited/8BEC59D02E2BB86A913CB97AAEB40F27.
Rabinowicz, Wlodek. 2017. “From Values to Probabilities.” Synthese 194 (10): 3901–29. https://researchonline.lse.ac.uk/id/eprint/66821/1/Rabinowicz_Values%20to%20probabilities_2016.pdf.
Raz, Joseph. (1975) 1999. Practical Reason and Norms. 2nd ed. Oxford University Press. https://academic.oup.com/book/9400.
Roussos, Joe. 2021. “Unawareness for Longtermists.” Presentation slides, 7th Oxford Workshop on Global Priorities Research, June 24. https://joeroussos.org/wp-content/uploads/2021/11/210624-Roussos-GPI-Unawareness-and-longtermism.pdf.
Scanlon, T. M. 2014. Being Realistic about Reasons. Oxford University Press. https://global.oup.com/academic/product/being-realistic-about-reasons-9780199678488.
Schuster, Carolin, Lisa Pinkowski, and Daniel Fischer. 2019. “Intra-Individual Value Change in Adulthood: A Systematic Literature Review of Longitudinal Studies Assessing Schwartz’s Value Orientations.” Zeitschrift für Psychologie 227 (1): 42–52. https://psycnet.apa.org/record/2019-18113-005.
Sengupta, Atanu, and Tapan Kumar Pal. 2000. “On Comparing Interval Numbers.” European Journal of Operational Research 127 (1): 28–43. https://www.sciencedirect.com/science/article/abs/pii/S0377221799003197.
Sher, Itai. 2019. “Comparative Value and the Weight of Reasons.” Economics and Philosophy 35 (1): 103–58. https://www.cambridge.org/core/journals/economics-and-philosophy/article/abs/comparative-value-and-the-weight-of-reasons/A3C3FFED0F7B88AF28EB68CFE2BEEBC5. Free version available at https://drive.google.com/file/d/1mBqVepp4_PqzMjyqqQQB3uNISCEKv14A/view.
Singer, Peter. 1972. “Famine, Affluence, and Morality.” Philosophy & Public Affairs 1 (3): 229–43. https://www.jstor.org/stable/2265052.
Singer, Peter. 1975. Animal Liberation: A New Ethics for Our Treatment of Animals. New York Review/Random House. https://archive.org/details/isbn_0394400968/page/n5/mode/2up.
Smallenbroek, Oscar, Adrian Stanciu, Regina Arant, and Klaus Boehnke. 2023. “Are Values Stable throughout Adulthood? Evidence from Two German Long-Term Panel Studies.” PLOS ONE 18 (11): e0289487. https://pmc.ncbi.nlm.nih.gov/articles/PMC10688669/.
Trammell, Philip. 2021. “Patient Philanthropy in an Impatient World.” Working paper. https://philiptrammell.com/static/Patient%20Philanthropy%20in%20an%20Impatient%20World.pdf.
Tucker, Chris. 2025. “Weighing Reasons.” In The Stanford Encyclopedia of Philosophy, edited by Edward N. Zalta and Uri Nodelman. https://plato.stanford.edu/archives/win2025/entries/weighing-reasons/.
Vaintrob, Lizka, and Owen Cotton-Barratt. 2025. “AI Tools for Existential Security.” Forethought. https://www.forethought.org/research/ai-tools-for-existential-security.
Vecchione, Michele, Shalom H. Schwartz, Guido Alessandri, Anna K. Döring, Valeria Castellani, and Maria Giovanna Caprara. 2016. “Stability and Change of Basic Personal Values in Early Adulthood: An 8-Year Longitudinal Study.” Journal of Research in Personality 63: 111–22. https://www.sciencedirect.com/science/article/abs/pii/S0092656616300502.
Wallace, R. Jay, and Benjamin Kiesewetter. 2024. “Practical Reason.” In The Stanford Encyclopedia of Philosophy, edited by Edward N. Zalta and Uri Nodelman. https://plato.stanford.edu/archives/fall2024/entries/practical-reason/.
However, as I argue below, the pessimistic induction from sign-flipping considerations does not bear evenly across strategies.
In terms of the response options described in the announcement post, I pursue option 1 of challenging a premise. In broad terms, the premise I critique is P2b about ‘best-guess’ comparison between two actions. But it is more precise to say that I challenge the maximality rule, as the narrow “always force” formulation of P2b is not my target. One can agree that “we shouldn’t always force ourselves to reach a ‘best-guess’ comparison” while rejecting the claim that a failure to satisfy maximality implies that we have no impartial reason to favor one option over another.
In decision theory, see Dietrich & List (2013), Dietrich & List (2017), and Cailloux & Meinard (2020). In moral philosophy, see Raz ([1975] 1999), Dancy (1993), Broome (2004), Horty (2012), Scanlon (2014), Markovits (2014), and Lord & Maguire (2016). See also the SEP entries on weighing reasons, practical reason, and reflective equilibrium.
For formal accounts on which the combined force of reasons varies incrementally, see Sher (2019) and Nair (2021). Views that involve incremental degrees of preference or betterness are explored in Fishburn (1970), Rabinowicz (2012), Knutsson (2021), and Hájek & Rabinowicz (2022).
On reasons as evidence, see Kearns & Star (2009) and Makins (2023).
See also the parallel between degrees of belief in epistemology and degrees of commensurability in axiology drawn by Hájek & Rabinowicz (2022, sec. 11).
On a larger range of value relations, see Hájek & Rabinowicz (2022).
Justified choice in the absence of full or determinate justification is discussed in Herlitz (2019).
Maximality has been formulated both in terms of preference and betterness. For example, Mogensen (2021, pp. 146–147) states maximality in terms of preference, while Kollin et al. (2025, pp. 9–10) state it in terms of ex ante betterness. I will liberally shift between both formulations, as my arguments apply to both.
I would have found it helpful if the sequence had clarified what assumptions license the move from maximality-based indeterminacy to the conclusion that we lack an overall impartial reason to favor one option over another. As a formal rule, maximality speaks to preference and rational permissibility rather than to the total balance of underlying reasons. Permissibility is generally a coarser notion than the balance of reasons: an option may be considered rationally permissible even if it is not the option we overall have most reason to choose. I return to this distinction in Appendix A.
The structural objections below are stylized, but deliberately so: if the maximality rule yields implausible verdicts even in cases designed to be easy, that gives us reason to doubt it in the hard cases we face in practice. Moreover, the practical objections that follow show that the problem is not confined to idealizations.
If the representor is required to be convex (roughly, if for any two distributions it contains, it must also contain every weighted average of them), the case can instead be constructed by taking P to be the convex hull of a set of extreme points, such that a single extreme point favors B by an arbitrarily small margin, while all other extreme points favor A by an overwhelming margin. The distributions depicted in the figure below can alternatively be understood as these extreme points. Since every distribution in P is then a weighted average of these extreme points, and expected-value differences average in the same proportions, no distribution in P favors B by more than the arbitrarily small margin of the extreme point that favors B.
In the formulation of the maximality rule found in DiGiovanni (2025c, sec. 3.1.1) and Mogensen (2021, pp. 147–148), the outlier would imply indeterminacy even if A and B were merely equal under the outlier distribution, which seems even less plausible. However, this implication can be avoided with a weaker version of the maximality rule that does not require the expected value to be strictly greater under each p ∈ P, as in Kollin et al. (2025, pp. 9–10).
We stipulate that A only impacts durations of torture in this thought experiment.
Moreover, according to the version of the maximality rule found in DiGiovanni (2025c, sec. 3.1.1) and Mogensen (2021, pp. 147–148), the preference would be indeterminate even if A’s imprecise expected impact ranged from preventing no torture to preventing all of it. And the range could again be made unbounded above, in which case the choice would essentially be between preventing 0 and [0, ∞) durations of torture. Maintaining complete indeterminacy in that case appears difficult to defend. Worse still, we can construct a case where, relative to the null action, A has expected impact [0, ∞) while B has expected impact (−∞, 0], and where the same admissible probability distribution p ∈ P assigns both actions an expected impact of 0. Standard maximality would also imply indeterminacy in this case, which seems especially difficult to defend. This challenge can also be extended to the weaker version of the maximality rule: say that A has expected impact [0, ∞) while B has expected impact (−∞, ε], and only one extreme point of P favors B, by the margin ε, while all other extreme points favor A. In this case, even the weaker rule would preserve indeterminacy merely because a single extreme point favors B by an arbitrarily small margin, despite A’s unbounded upside and B’s unbounded downside.
Others have also noted that this veto property is an unattractive feature of the maximality rule. For example, Mogensen & Thorstad (2022, p. 14): “[Decision rules such as the maximality rule] seem to give individual probability functions too much power by allowing them to veto any recommendation against choice of a given policy that would otherwise be made by one’s ‘credal committee’.” Lewis (2021): “a single member of one’s credal committee is enough to ‘veto’ a comparison of a versus a′. This doesn’t seem the right standard in cases of consequentialist cluelessness where one wants to make judgements like ‘on balance better in expectation’.”
To be precise, the claim is that verdict-flipping updates can be arbitrarily small, not that every arbitrarily small update flips the verdict. This is enough for the objection: wherever maximality places its threshold, some arbitrarily small update will be pivotal there.
In the words of Clifton (2025a): “I don’t think our reasons pin down particular intervals of numbers, either. The beliefs [about flowthrough effects] suggested by our reasons [are] much more like [Vague negative number, Vague positive number] than any definite interval.” DiGiovanni (2025a): “A representor still seems to be an imperfect representation of our epistemic state … The exact boundaries of these constraints … might be vague.”
Three distinct elements are captured in this illustration: the curve represents graded betterness, the vertical band represents imprecision, and the fading edges represent vagueness in the boundaries of that imprecision.
Whether we can reach unanimity will depend not only on how we assess and represent these unknowns, but also on how we understand “admissible credence functions.” These are non-obvious questions, and I do not attempt to settle them here.
Though one might ask whether biting this bullet would itself be favored by every admissible credence function. There is in fact a serious version of this concern at the meta-level of the rule’s own epistemic justification: if one finds maximality’s highly demanding standard of justification to be plausible, it is worth asking whether accepting maximality itself, given reasonable uncertainty and disagreement about the correct decision rule, meets a comparably demanding standard. If not, insistence on maximality would appear to involve a selective application of its own epistemic stringency.
Beyond these considerations about general plausibility, the maximality rule might also have significant strategic costs. For example, in strategic conflicts where a threatener can adjust the scale or credibility of a threat, the maximality rule may serve to raise the threshold that a successful threat must cross, thereby incentivizing more extreme threats or costly demonstrations. Maximality may invite maximal threats. This does not imply that maximality invariably worsens threats; for example, its demandingness may also deter weaker threats. But it may give us further reason to reconsider whether maximality is the most plausible decision rule to adopt, all things considered.
To be clear, I do not take a position on whether maximality is necessary even for full betterness. A more modest claim would be that maximality is merely sufficient for full betterness, while leaving open whether other, less demanding criteria may also be sufficient.
Γ-maximin favors the option with the highest minimum expected value across the credal set. This rule seems implausible since it would favor the null action in the earlier thought experiment in which a group of sentient beings stands to be tortured for life: the null action is guaranteed to prevent none of it, while action A’s imprecise expected impact ranges from adding a nanosecond of torture to preventing all of it.
I also take the general structural arguments for graded approaches over maximality to be the most important in practice. We usually do not fully write out expected-value calculations, and we will rarely need to specify an explicit graded measure. The central correction, in my view, is to move away from maximality’s binary mindset of unanimity or indeterminacy, and, more generally, away from the assumption that our comparative criteria must always take a categorical form.
However, both indices below could naturally be extended to the unbounded case by letting D(A) be 1 when sup(A) = ∞ and −1 when inf(A) = −∞, and by letting CD(A, B) be 1 when sup(A) = ∞ or inf(B) = −∞, and −1 when inf(A) = −∞ or sup(B) = ∞. For D, the index is left undefined when A is unbounded in both directions; for CD, the index is left undefined when A and B are unbounded in the same direction, or when either option is unbounded in both directions (though dominance may still discriminate in some such cases).
To be precise, we say that D is 1 when 0 ≤ inf(A) < sup(A) and −1 when inf(A) < sup(A) ≤ 0. For degenerate intervals [a, a], we say that D is 1 when a > 0, −1 when a < 0, and 0 when a = 0.
Since the length of the EV range above 0 is sup(A), while the length below 0 is 0 − inf(A) = −inf(A), the respective shares are sup(A) / (sup(A) − inf(A)) and −inf(A) / (sup(A) − inf(A)). The former minus the latter is (inf(A) + sup(A)) / (sup(A) − inf(A)). This can be rewritten as ((inf(A) + sup(A))/2) / ((sup(A) − inf(A))/2), which is the midpoint divided by the half-width.
While I independently thought of this simple index and its generalization presented below, it turns out that the generalized version is structurally equivalent to the acceptability index (also called the value judgement index) introduced by Sengupta & Pal (2000) in the interval-ranking literature.
This assumes, in each case, that the other endpoint stays bounded away from 0.
There will, under this graded approach, be a kind of shift where D goes from negative to positive (or vice versa). But, as long as the interval does not itself shrink toward [0, 0], this is a small and continuous shift in the midpoint tracked by small changes in D, not an abrupt jump. The only place there can be a discontinuous jump under this approach is at the degenerate interval [0, 0], and only along paths for which D does not itself converge to 0. Yet this is also the special case in which discontinuity seems theoretically plausible, since [0, 0] means that the two options are exactly equal in expectation under every admissible distribution. This is a determinate equality verdict rather than a mixed comparison that calls for a degree.
The general claim that a robust direct benefit shifts the graded verdict in the intervention’s favor does not require d to be the same under every p ∈ P, nor does it require the flowthrough range to be centered on 0. Suppose instead that the direct benefit is merely at least δ > 0 under every p ∈ P, and that the flowthrough range is [m − h, m + h] for any midpoint m. Then, when both the flowthrough range and the total EV range straddle 0, it follows that D(A) ≥ (δ + m) / h > m / h. In other words, it is enough that the direct effect is positive by at least some common margin across all models. This benefit shifts the overall assessment upward relative to the flowthrough effects considered on their own, even if the remaining uncertainty prevents a strongly favorable verdict.
This is not a form of bracketing as developed by Kollin et al. (2025), since the flowthrough effects remain part of the total evaluation and can affect both the degree and direction of betterness.
One proposal for extending the index D would be to make it sensitive to the broader profile of expected values across P. For example, a case in which a single distribution weakly dissents from an otherwise strong consensus in favor of A could then be treated differently from a case involving widespread and severe dissent, even when the two cases yield the same overall range. This might be achieved by supplementing the flat representor P with additional structure indicating which distributions are better supported, such as a nested family of plausibility regions or a non-probabilistic plausibility ordering. (That we can distinguish plausible probability functions from implausible ones suggests that some level of plausibility-discernment is possible, and it would be surprising if this capacity were restricted to exactly the binary ordering of plausible versus implausible.) I leave the development of such extensions as a project for further research.
The formula applies when h + h > 0; when both ranges are the same degenerate interval, we define CD(A, B) as 0. An alternative approach is to simply discard the dominance clause and instead rely on CD alone, truncated to the interval [−1, 1]. This may be more plausible if we think that what matters is the range of plausible expected values of an action, and not how those expected values happen to be paired with one another across the distributions in P. Note that the truncated CD reaches ±1 exactly when the two ranges are separated (i.e., when one range lies entirely above the other, with the endpoints at most touching), and separation in the strict sense entails pointwise dominance. The truncated index therefore agrees with the two-clause rule at the extremes, and differs from it only by assigning a graded verdict in the cases where pointwise dominance holds despite overlapping ranges.
In Clifton’s example involving b, we now go from a full preference for a over b to a preference for b over a to degree CD(b, a) = ε. Furthermore, in terms of ordering, pointwise dominance never conflicts with midpoint ordering but only breaks ties in it: if A dominates B, then inf(U(A)) ≥ inf(U(B)) and sup(U(A)) ≥ sup(U(B)), and hence M(U(A)) ≥ M(U(B)). It follows that the reversal from a full to a mild opposite preference is confined to a very narrow class of cases. An arbitrarily small sweetening can reverse a full preference only if M(U(A)) = M(U(B)), which together with the two inequalities forces inf(U(A)) = inf(U(B)) and sup(U(A)) = sup(U(B)). The reversal thus requires that A dominate B even though the two have exactly the same range of expected values. Once again, this is arguably the kind of special case in which a weaker discontinuity is plausible, since the discontinuity is in some sense found in the case itself: we suddenly go from pointwise dominance to non-dominance where the other option has a higher midpoint.
This is also the initial motivation presented by Clifton (2025b): “[Taking the midpoint] can be motivated by a symmetry intuition that I share to some extent.”
Arbitrariness can obviously be understood in various ways. Even if we specify it as the absence of defensible reasons, we still face the further problem of clarifying what it is to have defensible reasons for something. For example, one broad approach may be to understand the overall balance of reasons for a given position, and hence its degree of non-arbitrariness, in terms of our all-things-considered reflective equilibrium in light of all relevant considerations. This would be one way to understand the figure below. However, the figure below does not rely on a particular view of overall arbitrariness, nor does it assume that arbitrariness is literally quantifiable.
As earlier, we stipulate that A only impacts durations of torture in this thought experiment.
Some proponents of non-sharp credences likewise defend graded evaluations. For instance, Hájek & Smithson (2012) and Rabinowicz (2017) develop accounts that allow for non-sharp credences, while Hájek & Rabinowicz (2022) defend a graded treatment of evaluative comparisons. For related formal work combining imprecise probabilities with degrees of preference, see Montes et al. (2014) and Montes et al. (2017).
See, for instance, DiGiovanni (2025b, sec. 2.4, Q3) and DiGiovanni (2025d, sec. 4.1.3).
In calling these conditions relatively modest, I do not mean to imply that they are necessarily easy to satisfy. But compared to satisfying maximality in the realm of individual actions, these conditions can safely be called modest.
A strategy need not be the best possible strategy to be worth betting on. Since we can only choose among options we know of, it is enough that a strategy appears better justified than the relevant alternatives under consideration, or among the most justified when several are roughly tied.
Notably, many of the sequence’s objections to particular strategies can themselves be understood as identifying targets for further inquiry: they point to shortcomings in our current representations, such as a biased sample of hypotheses or an overly coarse representation of possibilities. Under a graded approach, these diagnoses can contribute to the case for targeted inquiry even when they fall short of establishing full betterness.
A useful analogy is the contrast between our intuitive physics and the modern science of physics that is informed by advanced mathematical models and superhuman amounts of data. Our comparative estimates may likewise become increasingly theoretically advanced and empirically informed, as they already have to some extent.
In one respect, this track record is a history of successful inquiry: the sign-flipping considerations were themselves uncovered by research, and often by rather modest amounts of it.
See, for instance, Dafoe et al. (2020), Clare (2023), and Ho et al. (2023).
By analogy, we can generally be more confident in a disciplined investment strategy than in any particular investment it includes.
The sequence divides its critique between high-footprint and low-footprint capacity-building. I focus on the latter, since that is where the strategies defended above mainly fall. I am also more tentative about high-footprint capacity-building, whose effects seem harder to anticipate and control.
UEV stands for unawareness-inclusive expected value.
As mentioned earlier, an option may be considered rationally permissible even if it is not the option we have most reason to choose.
Even if one thought that the maximality rule is by far the best formal decision rule we have, it would still not follow that it fully captures our impartial reasons.
On the stability of personal values, see Milfont et al. (2016), Vecchione et al. (2016), Schuster et al. (2019), and Smallenbroek et al. (2023). On the continuity of prosocial dispositions, see Eisenberg et al. (2002). There is also evidence that, on average, “political attitudes are remarkably stable over the long term” (Peterson et al. 2020, p. 600). See also Evans & Neundorf (2020).
For example, Peter Singer has been writing about altruistic donation and animal suffering for more than five decades; Frances Power Cobbe campaigned for social reform and against vivisection for several decades; and Abdul Sattar Edhi spent 65 years building Pakistan’s largest network of orphanages and shelters.
See, for instance, Matsumoto et al. (2016), Vecchione et al. (2016), Smallenbroek et al. (2023), Pollerhoff et al. (2024), and Li et al. (2024). On the continued development of moral judgment and value reasoning in adulthood, see Colby et al. (1983), Armon & Dawson (1997), and Armon & Dawson (2003).
The first challenge is that the wager does not justify the status quo, since there are still many open questions given precise subjective Bayesianism (PSB), including how to handle unawareness. Yet this fair critique of the status quo does not seem to speak against PSB itself. Another challenge concerns inter-theoretic comparisons of different moral views: if we have qualitatively different moral frameworks, “it’s not clear how we can say, ‘The moral weight of (i) your expected impartial impact from the PSB perspective overwhelms the moral weight of (ii) e.g., your parochial or non-consequentialist reasons.’” But this challenge seems to apply equally to views that combine imprecise credences with maximality, whether in the context of wagering or otherwise. So this challenge does not seem to speak against wagering on precise Bayesianism in particular. Finally, there is a claim that PSB embraces arbitrariness. Yet it is unclear whether PSB is all-things-considered more or less arbitrary than imprecise credences combined with maximality. That seems like an open question, and the answer may depend on which specific premises and desiderata one finds more or less plausible.
Such partial convergence is not surprising, considering that the wagers often end up reflecting the same underlying reasons. For instance, both precise credences and graded approaches are sensitive to mild sweetening and will track the same underlying evidence with similar directional updates. Likewise, the rationale for doing research under the moral-reflection wager will also have substantial force under the other wagers, namely to improve our practical verdicts under an outcome-focused framework. This is highly useful under each wager, even if to a different degree across them. Furthermore, that many views and wagers may converge to endorse foundational research is not so surprising considering that the future may contain sentient beings for billions of years, while systematic exploration of these questions is only a few generations old. In other words, considering our limited state of knowledge and where we are in time, the explore-exploit tradeoff likely tends to favor exploration across a wide range of outcome-focused views and wagers. That said, this shared grounding does somewhat limit how much additional support the convergence provides: agreement among wagers that draw on overlapping reasons is not independent confirmation in the way that agreement among independent lines of evidence would be. But it still supplies a meaningful form of robustness, namely robustness to our uncertainty about which framework or starting point is correct. A recommendation that survives across several plausible starting points is less likely to be an artifact of any one of them. By analogy, when several research teams analyze the same dataset using different modeling choices and reach the same conclusion, this provides stronger support than a single team’s analysis, even though no new data has been gathered.