Epistemic status: This post reflects my own reasoning about a methodological question in cost-effectiveness analysis. I'm not a professional economist or a member of any EA research organization, and I'd welcome pushback from people with more technical expertise in this area.
Note on AI use: A significant portion of this post's text was drafted with the help of an AI language model based on my outline and ideas. I've reviewed and edited it for accuracy and tone before posting.
Summary: Cost-effectiveness estimates (like "$X per life saved" or "$Y per DALY averted") are central to how the EA community allocates resources, but a single point estimate can create false confidence. This post argues that we should treat these numbers as ranges shaped by real uncertainty, not fixed facts, and suggests some practical ways to hold that uncertainty without becoming paralyzed by it.
A lot of decision-making in effective altruism, from individual donation choices to large grantmaking decisions, leans on cost-effectiveness estimates. These numbers are genuinely useful. They force explicit reasoning, they're comparable across interventions, and they push us toward transparency about our assumptions. But there's a risk in how we use them: a single number, reported without its surrounding uncertainty, can look more solid than it actually is.
When an estimate says an intervention costs a certain amount per life saved, it's easy to treat that figure as settled. In practice, it's usually the output of a chain of assumptions, each with its own margin of error, stacked on top of each other.
A few sources of uncertainty tend to get compressed out of a final point estimate:
Underlying data quality. Many interventions are evaluated using data from specific regions, time periods, or implementation partners. Generalizing that data to a different context, a different country, a different decade, involves an implicit judgment call about how similar the two settings really are.
Modeling assumptions. Converting a measured outcome (say, a reduction in disease incidence) into a standardized unit like a DALY or a "life saved equivalent" requires assumptions about things like discount rates, age-weighting, and how to handle indirect effects. Reasonable, well-informed people can and do disagree on these choices, and different reasonable choices can shift a final estimate meaningfully.
Publication and selection effects. Interventions that get rigorously evaluated aren't a random sample of all possible interventions. This doesn't mean evaluated interventions are bad, but it does mean we should be cautious about assuming an intervention that hasn't been studied is worse just because it lacks a number.
Diminishing or shifting returns. A point estimate is often calculated at a particular scale of funding. An intervention that looks highly cost-effective at a small scale may look different once significantly more money is directed toward it, since the easiest, cheapest opportunities tend to get taken first.
None of this means the estimates are worthless. It means the single number at the end is better understood as the center of a distribution than as a fact.
One low-cost habit that could help is reporting cost-effectiveness figures as ranges rather than single numbers whenever the underlying analysis supports it, for example "$X to $Y per DALY averted, with a central estimate around $Z," rather than just "$Z per DALY averted."
This isn't a new idea, some organizations already do a version of this, and sensitivity analysis is a standard part of rigorous cost-effectiveness work. But even where the underlying sensitivity analysis exists, the headline number that gets repeated in donor-facing summaries, forum posts, or casual conversation often drops the range and keeps only the point estimate. That compression is where a lot of the false confidence creeps in, not in the original analysis, but in how it gets communicated afterward.
I don't think the right response to this uncertainty is decision paralysis. A wide but well-reasoned range still tells you a lot, especially when comparing interventions that differ by an order of magnitude or more. If one intervention's plausible range is $50 to $150 per unit of impact and another's is $3,000 to $8,000, the ranges don't need to be precise for the comparison to be informative.
Where I think the uncertainty matters more is in comparisons between interventions whose estimated ranges overlap significantly. In those cases, treating the point estimates as decisive, rather than roughly tied, risks over-updating on differences that may not be real.
I'm genuinely unsure how much this concern should change actual behavior within the community. It's possible that people doing this kind of analysis professionally already internalize this uncertainty well, and that the issue is mostly about how findings get summarized for a broader audience rather than a flaw in the underlying work. I'd be interested in hearing from people closer to this work whether that matches their experience, or whether they think point estimates get overweighted even among specialists.
Cost-effectiveness numbers are one of the more valuable tools this community has for reasoning about impact, but a single number at the end of a long chain of assumptions deserves to be held a little more loosely than it often is. Reporting and discussing ranges, rather than just point estimates, seems like a small, low-cost change that could make our collective reasoning slightly more honest about what we actually know.