"It is not even wrong" - Wolfgang Pauli
Summary: I identify three IMO fatal flaws of the formal ITN framework. 1) It looks like it proves the ITN heuristic, but it doesn't. 2) It looks like it proves a specific claim about ITN (10x higher neglectedness => 10x better to work on) but it doesn't. 3) It is not better than the alternatives (typically). I also offer two other contributions: A) A satirical model which I think demonstrates what's wrong with formal ITN ("The Geese Aggression Index") and B) I point to specific instances where mentions of the model need to be updated or removed.
The Importance, Tractability and Neglectedness framework (henceforth: ITN, the factors are also known as: Scale, Solvability and Crowdedness) is a heuristic used to identify and prioritize cause areas. The main idea is that bigger and more neglected problems may be more cost-effective to tackle. This framework has been widely used in EA and has been adopted by 80,000 Hours, who have helped develop[1] a formal version of this framework. Based on what I've read, they don't seem to actually apply the formal framework to create their rankings (I'm unsure if they have done so in the past though). They do however have this formal version in their website and seem to use it as evidence that "roughly speaking, the three factors multiply together" (a claim that made it into their career guide!). Here is the model in 80K's website:
If you multiply the factors together, we get back to welfare/additional resources:
Now, I want to make something clear: this derivation is mathematically correct (under all circumstances). I believe it is misleading, and I don't think it is very useful. But it is technically correct.[2]
Now both the formal model and the heuristic have been subject to quite a bit of critique by the community. A review of the critiques was made by DirectedEvolution 7 years ago. I would like to point out that I have found that I agree with many of the critiques that have been raised and personally I think that the community (and 80K) have not updated enough because of them. This post is somewhat similar to a post disclosing Founder's Pledge own critique and improvements of the model and I also expect to echo other points made by the community.[3] However, I believe there is value in this post. I think many are unfamiliar with the arguments I will present today, and I also point to specific instances where mentions of model need to be removed or updated.
It is very useful to take a look at the definitions of the factors first. Let's start with scale and take the cause area of Animal Farming as an example. What is its scale? Outside of the formal framework I would probably consider the number of animals in the whole industry. But a better definition might maybe be how many years of suffering the industry causes, which is what's given by the model:
Scale := Good done / %of a problem solved
So, scale essentially answers, by solving 100% of the problem, how much welfare do we get. This does seem roughly like what we meant. What about solvability?
Solvability := % of a problem solved / % increase in resources
Now, I'm not sure about you, but I don't think I ever really had a definition for solvability before engaging with this framework. It always made intuitive sense to me that some problems are just harder to solve than others. But I never thought about the meaning of solvability. I never spared a thought about how it could be quantified or how it could be measured.
Notice that using solvability as a part of the heuristic makes sense: Say I want to propose converting Mars into a gigantic animal sanctuary for happy rabbits. Scale and neglectedness will be on my side, but solvability will serve as a reality check. However, if we try to use numeric values for solvability (instead just of "low" and "high") we might then realize that we never had a definition for it (at least I didn't). And perhaps more interestingly, the first thing that may come to mind is cost-effectiveness, which is what we were after in the first place! This has been pointed out before and it has been argued that we should rely exclusively on cost-effectiveness instead of ITN (and also that we should do Fermi estimates of it when necessary).
Note also that solvability is defined here as a ratio where the denominator is a percentage increase, what's called semielasticity[4]. On the other hand, cost-effectiveness is defined as a simple ratio[5]: problem solved / increase in resources. So, the formal model uses something like cost-effectiveness, but that is not really cost-effectiveness. This might seem like something petty to focus on. But mathematics is ruthless, and if the definitions don't match, the formal model can offer almost no evidence for the heuristic.
Lastly, let's look at neglectedness:
Neglectedness:= % increase in resources / extra resource
This definition does track with intuition. If we want to add 1 unit[6] of extra resources, then for neglectedness we get:
which = 1/(current resources), which might be what we mean by neglectedness (maybe).
However, this should be surprising! This is a mathematical formula that shows that neglectedness is a factor in welfare / dollar (and it is correct!) And since it is multiplying, we should believe that more neglected cause areas *must* have a higher good/dollar. But the idea that neglectedness should matter at all comes from the concept of diminishing marginal returns, which is an assumption. It is not a law of the universe. So why does it show up on the formula?
The 80000Hours' website and career guide claim the following:
(I am going to call this "the 10x claim" from now on). To assess the valitidty of this claim, let me present you with three statements regarding the formula. Here it is again for reference:
Statements:
Statement 1 is true. Statement 2 is wrong: the formula doesn't give us enough information to make that call. And statement 3 is generally[7] not true. To see why remember first that earlier we derived that neglectedness is just 1/resources.
Then, notice that solvability's denominator, "%Increase in Resources", can be broken down into additional resources / current resources.
Rewriting reveals that...
... we first divide by neglectedness and then multiply by it! So neglectedness appearing in the formula is kind of artificial. Indeed, the formula is compatible with neglectedness having no impact at all!
To make my point clearer, let's assess two different naive ways we may use this formula:
Now, with the silliest assumption, the claim that 10x higher neglectedness => 10x welfare/resource is obviously true. The only problem is that the assumption is a mathematical impossibility.
Conversely, with the silly assumption, the claim that 10x neglectedness => 10x welfare/resource is obviously false. Any change in neglectedness is canceled out by solvability.
The problem is that it is very easy to look at this model and think that "the silliest assumption" applies and thus the model proves the 10x claim. Yet this is simply not true. It looks like this is the case because the formula is given as a multiplication of 3 different factors, so the idea of increasing one while leaving the others constant comes naturally.
Additionally, while I don't think "the silly assumption" applies, notice 80K's website does not provide us with evidence it doesn't. The model is completely compatible with neglectedness having no effect whatsoever. So how do we actually get the 10x result? Here's how:
The 10x assumption: assume logarithmic returns.
Mathematically, the only[8] way to get the model to spit out the 10x claim is by assuming logarithmic returns to resources (see proof in the "I use math for your reading pleasure" section). Let me explain what this means:
The model presupposes that "Good Done" is a function of "Problem Solved" which is a function of "Resources". So, it looks like this:
Then it is also reasonable to define "Good Done" as a function of resources:
And we can prove that this function must be of the form:
(ln is the natural logarithm, k and C are constants)
And with an additional reasonable assumption[9] you can go a little further and prove:
And it's just not clear why we should think that's the case. 80,000 Hours gives no evidence for this. At the community level we have seen some arguments for logarithmic returns from Owen Cotton-Barratt, but he himself admits that for many (most?) cause areas this is just an ignorance prior[10]. I am unaware of any other evidence for this assumption.[11]
The claim is not completely ludicrous, but it's not completely justified either. So, the evidence for 80K's claim is that if you pick the right assumptions the factors roughly multiply together. You kind of have to assume the conclusion.
It is true that it makes sense to assume that with more resources, we will get somewhat worse at tackling a given problem. After all, we expect the low hanging fruit to be picked first. This is called diminishing returns. Logarithmic returns is one specific case of this. But it is not the only way to make this assumption: square root returns, for example, also fit this criterium. And if we asume:
then multiplying neglectedness by x would result in sqrt(x) higher cost effectiveness. For high differences (>100x) in neglectedness between cause areas, root returns predict an effect on cost effectiveness orders of magnitude smaller than log returns. So, the specific function does matter.
I've stated before that I think the community has been slow to discard the model even with mounting evidence against it. It is because of this that I'll allow myself to do something a little impolite: mock the model.
Here I will present a "breakdown" of cost-effectiveness that is mathematically correct, yet absolutely ridiculous. Notice that it has a very ITN flavor to it. I hope it serves as proof that true =/= useful.
Here is the bumper sticker version ("Progress" is a renaming of "Problem Solved"):
This cancels out and we get back to Good Done / Resources:
(There is a mathsy version of this in the next section)
I have now "proven" that good/resource is a function of scale, the mysterious factor X and the Geese Aggression Index™. I may now claim that cause areas with a 10x higher Geese Aggression Index[13] are, all things equal, 10x better to work on. This will ignore, however, that "all things" cannot possibly be equal because X will necessarily change to counteract the effect of the Geese Aggression Index.
In this section I show the same cancellation from before with more mathsy notation (skip if you don't like math) and then show the proof of log returns I promised. However, I must make an awkward point first: the two percentages in this model mean different things (Yes this is confusing. I'm sorry, this is not my model. Understanding this is not vital.).
One is a "percentage of" and another is a "percentage increase". It is helpful to think of an example. Using the Animal Farming example, we could think of freeing an additional chicken out of a cage, when there is a total of 100 chickens in planet earth. In this case, "% of problem solved" means that we freed 1/100 chickens, so 1%. In the second case we can think of donating an additional dollar to the cause (let's say people have donated 5 dollars total up to this point). Then it doesn't matter if it would take 100 dollars to solve the problem, we calculate the percentage relative to the 5 dollars donated already, which is 1 / 5 = 20%. This is what's meant by "% increase".
So, in the first case 100% means completely solving the problem and in the second case it means doubling our resources. I know this is what 80K means because it is consistent with how they talk about the model.[12]
Math:
And now the proof of log returns:
And the "mathsy" version of the Geese Aggression Index:
As I said: 80000Hours' career guide claims the following:
But this is a very shady claim. For starters: the definitions in the formal version differ from intuition (for solvability mostly). If the evidence you have for something is "given my preferred and unusual definitions, the claim must be true" maybe you don't have much evidence. But it gets worse, because "holding everything else equal" doesn't make sense in this formula (as per my arguments above).
Outside of this model we do have a some evidence of why ITN might multiply together (log returns), but they don't seem to base their claims on that. Thus, I would honestly recommend that 80K removes this claim from their career guide, or that they update it to reflect where this assumption comes from.
I originally was going to say that 80K never claims this is a "proof" of ITN. While writing the section above I noticed that's not true. Still, were 80K to update their career guide to no longer say this I would still fear this would seem like a "proof" of ITN. Therefore, I think that the formal model website's needs to be updated in some way that does not give the impression that this is a proof of ITN. If they want to keep the model in their website that's great, but a disclaimer would be nice.
I should also note that this model has been criticized for being misleading in a different way. It has been argued that it helps obscure the actual reasons / calculations we think something might be worth doing by instead using a more socially acceptable formula.
ITN as a heuristic is helpful. I think good arguments have been made that it is flawed but I would still consider it an achievement of the EA community. This is not true of the formal model. The whole idea is to have a way to estimate cost-effectiveness. But, as it has been pointed out before, we always implicitly estimate cost-effectiveness when we estimate solvability. Here's why:[14]
Solvability is then cost-effectiveness multiplied by something else. The rest of the formula simply divides by that something to get back to cost-effectiveness.[15] It is simply not clear why estimating cost-effectiveness[16] in its semielasticity version should be better than the usual version. Indeed, instead of just estimating (guesstimating?) one factor, we estimate three. Seems less accurate.
This is a bit more of a subjective point. But the whole reason why I am writing this post is because I wanted to use the quantitative version of the model to actually make a career choice (if that sounds weird maybe that is because it is. I did think it was probably not a very practical thing to do but I decided to give it a shot anyway). It seems now clear to me that the correct way to do this is to try to Fermi estimate cost-effectiveness (and maybe use scale and neglectedness as informing priors on how cost-effective the average intervention in the area might be). So engaging with the framework in the first place was probably a mistake on my side. I wonder if it could also be distracting for other people, as I think the model itself is not very helpful.
"All models are wrong, but some are useful" - George E. P. Box and Norman Draper
If you want to keep a model for didactic reasons you can definitively do so. Here is one such model by yours truly (A similar model [17]has been proposed by Founders Pledge):
From there it's easy to suggest that less resources and higher scale lead to higher cost-effectiveness (slope) while also being transparent that this is basically an assumption. Here is one instance of the same model with the assumption broken:
I also think you can grab this idea and build richer models that include other things, like, for example, personal fit. I think there may be some value in doing this, but I'll abstain from doing it in this post.
Calculating cost-effectiveness directly is always the best option. If necessary do a Fermi estimate (see here for examples of Fermi estimates of cost-effectiveness being used successfully within EA). But I can also imagine that building cookie-cutter cause prioritization models (like this one) could be useful. Maybe even models based on ITN.[18] I plan to explore that in the future.
For all its faults, I actually kinda like the model. I think there is an advantage in thinking about cause area prioritization quantitatively and perhaps that needs more exploring. I think there is value in thinking about ITN more deeply. And I think the model can be a starting point to designing models that are more transparent and useful. However, I still think that it must go.
The formal ITN model fails to be what it should. It is neither a proof of ITN as a heuristic, nor does it demonstrably help in estimation of cost-effectiveness. It is also untransparent and is particularly misleading when it comes to neglectedness. There are alternative models for didactical purposes (you can make your own, it's easy) and there are alternative ways to estimate cost-effectiveness. Despite this, the model has found its way to places it shouldn't have, e.g. factual claims made in 80K's career guide.
Oftentimes, when people introduce the model the say stuff like: "this model is true and mathematically trivial. We are not sure if it's useful though." I hope I have now successfully answered that question: it's not.[19] And though many of model's faults have already been made clear to the community, we have yet to fully throw it out. It must go.
This should shift the discussion about cause are prioritization. First: cost-effectiveness is the queen of all considerations (so bringing up neglectedness after cost-effectiveness has been measured is useless). Second: if we are going to argue for/against a cause area based on ITN, we should not use the formal definitions (1/resources is probably a suboptimal definition of neglectedness; If you argue something is neglected because people find it boring or weird, you have a stronger case). Third: the jury is out on any other factors that may correlate with cost-effectiveness. Feel free to theorize about them, or better yet, measure them!
I'm grateful to @Johannes Riemenschneider, for providing ample feedback to an early draft. All errors are of course mine!
I'm actually not entirely sure it was 80K who developed it. In their website they mention Coefficient Giving being the ones who came up with "the framework". By that I think they mean the heuristic. If that's the case, then the formal model seems to have been the result of a collaboration of 80K with the staff of the Future of Humanity Institute.
The best kind of correct.
In all likelihood there have been people before me that have perfectly understood all of the faults I point out here.
Indeed, fellow economists, mathematicians and nerds in general. In this specific version of the model, the solvability term is not an elasticity, even if it looks that way. This is because the two percentages in the model are not the same kind of percent (this is what happens when you describe a model in words and not formulas). See the "I use math for your reading pleasure" section for an explanation. But in short, we have a semielasticity, as the first percentage is just dividing by a constant, but the second one is indeed about proportional change. (I know this is what 80K means because it is consistent with how they talk about the model)
There exist, however, versions of this model where this term is indeed an elasticity (and the term for scale is thus also different).
(technically a derivative)
Notice that my career choice between, say, AI Governance and Animal Welfare shouldn't depend on whether I plan to dedicate 40 or 20 hours to them. So, we can set the contribution to any value. Choosing 1 is very convenient.
In the mathematical sense of the word "generally".
This is not exactly true. This is true if at every level of resources, the 10x claim holds exactly. I am also assuming, as 80K asks to, that everything else stays equal. But you get the point.
By assuming scale/importance is constant. Example: we assume that every additional chicken we get out of a cage is worth just as much as the last one.
But he thinks for problems where success is made by a bunch of small contributions this is theoretically justified. A good example of such a field is research. And studies have indeed shown that research returns are indeed logarithmic. It just doesn't apply to *every* field
But I'd be really excited to see people try to test it with real EA cause areas data!!
See how they say we should assess scale and solvability.
Which is ofc defined as how aggressive geese are on average towards a person working in a given cause area.
I should probably get familiar with LaTeX.
This is not exactly correct, and it depends on your definitions. Scale does actually provide some useful info, and I could even concieve of tweaking the definition of neglectedness to capture useful things, like personal fit. But this is beyond the point.
I am not sure if I'm mixing up cost-effectiveness and cost-efficiency here. But you get the point
You can find a full explanation on page 35 of this report.
Consider this: (in my notation) dG/dr = dG/dP * dP/dR * dR/dr. Moral importance per unit * cost-efficiency * personal fit. It's the same model but without the neglectedness/solvability problem(s) and the last term becomes personal fit. Here you assume R = Rc + PersonalFit * r.
Or rather: I hope I have now convinced you of what the community had already successfully pointed out much before me.