This competition entry has been selected for publication by the Forum team.
Our uncertainty about the future is vast. Unfathomably vast.
Imagine that only 1 in a million people on the planet will make decisions that impact the future, and that these 8000 people will only make 1 such yes/no decision in their lifetimes. With this very simple model, the number of potential futures is 2^8000 or approximately 10^2408. And that’s just for the people currently alive. If the number of future humans grows by orders of magnitude, then the number of potential futures would grow absurdly large. Considering that experts can disagree by over ten orders of magnitude over short-term future probabilities, any charting of a course over the long term future seems doomed.
Given this vast ocean of uncertainty, it seems that trying to affect effective long term change is completely hopeless. This is the premise of Anthonny Giovannis series on cluelessness, currently the subject of an EA forum competition. These articles make the argument that we are completely clueless as to the best course of action for the long term future, or after a technological singularity.
My main issue with these articles is that they are built up on too narrow of an epistemological foundation. For one, a large amount of it is specific to bayesian epistemology and Rationalism, with a significant portion of his further resources pointing to these sources. As a non-bayesian anti-Rationalist, this puts me in a different universe to the author. Even further, the author posits some very specific forms of bayesianism, such as “unawareness-inclusive expected value” and indeterminate beliefs. One might get the impression that merely rejecting these formulations makes the problem of cluelessness go away.
Despite all this, I think the underlying arguments are quite good. In this article, I will try to put forward a different framework using most of the same arguments that is hopefully more robust to ones underlying epistemology.
One other issue with the articles is that it takes long-termism as a given and doesn’t discuss short-termism much. It says we are clueless about changing the future a million years from now, and are not clueless about whether eating a burger will alleviate our hunger, but it doesn’t discuss much about whether or not we are clueless about whether malaria vaccines will help impoverished people (I believe we are not clueless about this).
To sort this out, I thought of an interesting and kinda silly analogy that I think could help elucidate some of the issues here, by likening the oceans of possible futures with the actual ocean. It is about a man named Doug who does not know what a fish is. It was heavily inspired by figure 4 in this cluelessness article.
I will explain several variations on the analogy, and explain why they provide good defences for short-termism, but not for longtermism. I don’t pretend to be a philosophical expert or anything, and I want to be clear that I did not invent most of the arguments in this article: many of them are altered or adapted from the arguments in this cluelessness article and in other articles I have read over the years, such as the great work from David Thorstadt.
Doug is a man who in some weird twist of fate, has no idea what a fish is. He has never seen any fish in person or in pictures. One day his coworkers start talking about fish, and he has no idea what they are talking about. They explain that fish are creatures that swim around in the ocean. He is shocked to hear that there are unimaginable numbers of fish in the ocean.
Before the conversation can continue, a coworker excitedly remembers that he has a friend Carol that lives nearby with a fishtank. They take Doug over, where he looks at fish for the first time. In the tank, he counts 4 green fish and 1 red fish. He is suitably impressed at the sight.
Later on, he is chatting with friends about this new revelation, and idly wonders about what the rest of the fish in the world are like. His friend asks him to guess: they ask him to estimate the probability that there are more green fish than red fish on the entire planet.
If forced to bet, what odds should he answer?
Of the fish that Doug knows about, 80% are green and 20% are red. Does that mean that he should guess that roughly 80% of the fish in the entire ocean are green, and hence it is very likely green fish outnumber red fish?
Of course not! The contents of the fish tank are not at all a random sample of the fish population. The prevalence of the green fish in the tank could be due to any number of factors, unrelated to the prevalence of fish in the entire ocean:
For these reasons, the prevalence of green fish in the fish tank tells us basically nothing about the prevalence of green fish in the ocean at large. Therefore if Doug was in a state of cluelessness about the ratio of fish colour before he saw the tank, he should remain in a state of cluelessness afterwards.
An interesting thing about this analogy is that with some small tweaks, you can get alternative versions where you actually be pretty sure about the superiority of green fish. I will introduce these alternative versions later, as they will form the backbone of various defences against cluelessness.
For now, let’s move to a situation which I believe is analogous, involving a longtermist effective altruist:
Dave is an effective altruist who has recently sold a tech startup for a large amount of money. He wants to use his money and influence to make a meaningful difference to improve the long term future of humanity over the next million years. In order to play to his strengths, he is considering starting up his own safety-focused frontier AI development lab.
To estimate the effect of his actions, he hires a think tank to try and estimate the long-term effect of his actions on the future welfare of humanity. The think tank spends some time looking over the problem, and comes back to him with some results.
They say they have explored five different plausible scenarios of equal likelihood resulting from founding the AI lab. Four of them have net positive effects on the future of humanity, while only one has a negative effect. They plug these values into an expected value calculator and claim that the intervention will in expectation improve the lives of tens of trillions of people.
How confident should Dave be that the intervention is actually net positive for humanity over the long term?
Of the future scenarios presented to Dave by the think tank, 80% are positive expected impact (which we will colour green) and 20% are negative expected impact ( which we will colour red). Does that mean that he should expect an actual 80% chance of positive impact among all possible futures in the real world, and invest accordingly?
I believe that the answer is no. In this scenario, the ocean of possible future impacts of a major decision, extending out a million years in the future, is so gargantuan in number that it makes the number of fish in the ocean look miniscule in comparison. When the think tank claims to have investigated five scenarios, they might have looked at many possible outcomes in significant depth and effort. The play models inside a think tank are like the play fish inside the fish tank: a tiny, non-representative sample of all the vast ocean in the outside world.
If anything, the situation for Dave is worse than the situation for Doug, because at least Doug gets to see some actual fish. In contrast, Dave never sees the actual long term future, only current day guesses about it. Our guesses about the future are not the actual future: If we want to make decisions about the latter based on the former, we have to prove that there is an actual strong link between them, despite the massive uncertainty about the state of the future and our fallibility as human (or AI) estimators.
And there are plenty of reasons that a think tank could declare the proposed AI startup to have positive impact, even when this is not true in reality:
Anthony gives some other reasons to doubt the unbiased nature of estimators here.
If we follow this logic to it’s conclusion, then Dave should end up in the same situation as Doug: If he was clueless about the impact of his intervention before commissioning the think tank, he should remain clueless after. If we believe that we should initially have no idea whether or not our actions will help or hurt the long term future, then this should be unaffected by the think tank research, even if they put a lot of time and effort into it.
When making analogies and arguments like the one I made above, you have to be very careful about proving too much. After all, we could make similar arguments about short term interventions like distributing malaria vaccines to impoverished villages, or take it to the extreme and say that there is no reason to save a child drowning in front of you because you can’t predict all the possible consequences of your action.
I am no moral nihilist, and I believe that helping people is good. In order to properly dismiss longtermism on cluelessness grounds, I must defend everything else from the same arguments.
I think there are several reasons why the analogy above is correct for talking about million-year scale longtermism, but is not correct for short term stuff like malaria vaccines or saving a kid from drowning.
To explore these reasons, I will return to our original fish analogy, and make mild tweaks that dispel the cluelessness feeling. I will show how each tweak can be turned into defences for short-termist interventions, but are comparatively weak for defending long-termist ones.
Suppose instead of being asked to estimate whether there were more green or red fish in the entire ocean, Doug was asked to estimate whether or not there were more red or green fish in the fish tanks of students in his university, where he is informed there are like 20 fish total.
In this case, the fish in the tank make up a meaningful portion of the actual thing we are measuring. If each remaining fish had an exactly equal chance of being red and green, the expected proportion of green fish to red fish would be around 56%, and the chance of green being higher overall would be around 75%.
The analogy to decisionmaking here is in cases where we expect our scenario prediction to be accurate and to also cover a large portion of the relevant possibility space. For example, when you go get a burger, you have a pretty good idea about the likely outcomes of your events (eating a burger). You may not have mapped out every possible scenario, like getting hit by an asteroid in your car on the way, but these events are extremely unlikely and so do not take up a large portion of the future if you weight by probability.
This defence will apply if you have a very strong reason to expect your guesses to successfully represent the future due to empirical evidence about the problem at hand. This could be due to scientific evidence, past record of successful predictions on similar problems, or an extension of scientific knowledge such as the laws of physics.
It will also apply to situations where the scope is very short term, like the “drowning child” scenario. If your goal is to save the lives of children in the short-term, and you see a kid drowning in front of you, you should save the kid.
In all these situations, you might have missed something relevant, but it is unlikely to be a large enough miss to cancel out the obvious good that is going to result from your actions.
This defence relies on the “ocean” of possibilities being small and predictable, so you can accurately capture a large enough portion of them in your reasoning. There is simply no defence of this sort for longtermism. No matter how hard you work, there is no way you can sift through all the possible futures out there extending out for hundreds or thousands of years.
Imagine that we modified the fish analogy so that instead of the fish tank being curated by some random student, it was instead the result of a global fish sampling project. This project created a process to select a sample of fish, completely and perfectly randomly, from all around the world.
This change would drastically change the conclusions we took from the story. If the fish in the aquarium were a random sample selected from the entire ocean, it would imply a 97% chance that there were more green fish than red fish.
This is the result of random sampling, the statistical principle that make polling work. You don’t need to poll every person in the country to get a good indication of how people are going to vote, because if you are randomly sampling, the statistics of your sample will rapidly approach the statistics of the population as a whole. The analogy I’ve seen is that you don’t need to taste the whole soup to know how it tastes, a spoonful is enough if the soup is well mixed. Sampling is never truly random, but it can be good enough for many purposes.
The defence here is the claim that more likely future events are more likely to end up in our hypothetical reasoning, in proportion to their likelihood. For this to be true, we have to have strong empirical reasons to believe that we have strong predictive skill with regards to whatever scope we are studying.
I can think of a circumstance where this applies especially strongly to short term prediction: cases where our prediction of the future is an extension of a smaller scale experiment, such as a randomized control trial. Say we are trying to predict the outcome of rolling out a malaria vaccination campaign to a million people in impoverished parts of Africa. We do a trial run in ten different villages, , and build up solid scientific evidence. You then predict that the results of the full run will be the same as in the villages you study. Although you haven’t seen the future, and you have only evaluated the vaccine in a small portion of the population, you can still predict that the results of the full trial will be good. This is taking advantage of the random sampling and assumed similarities between humans from different places: you would not expect the same number of lives to be saved if you deployed the vaccine in a first world country, for example.
Can a similar defence be applied to Doug and his AI company? No, not really. None of the beliefs or predictions about the future are based on scientific trials, and there are not really any regularities between current day events and future ones to take advantage of here. The events they are trying to forecast are vast and world-changing, with no small scale experiments that can be scaled up to give us anything. There is no empirical track record of predictive success on a long-termist scale, and thus no reason to expect that we are randomly sampling from the future.
Suppose we revert back to our initial fish analogy, but right at the end, Doug hears a snippet from a nature documentary, saying “red scales are significantly more energetically costly to grow than green scales”.
In this case, Doug should probably guess that green is more common than red fish. The reason is that since red scales are more costly than green scales, we could reason that more species of fish would evolve green scales, and that green fish were overall more likely to survive. This is by no means certain reasoning (perhaps red scales confer huge advantages in other areas), but it is sufficient to give better odds to green fish.
Crucially, the belief in better odds to green fish has almost nothing to do with the proportion of fish in the tank. He should give the same answer if the fish proportion was the other way, or if he had never seen the fish at all.
If you have a strong reason, ahead of time, before doing any modelling of the future, to believe that an action would help, then this defence applies.
For shortermist equivalent, i’d point to any situation where it’s blindingly obvious that the effect of an action will be helpful, without needing to do any serious modelling. For example, if a bomb is about to go off in a crowded area, is disarming it a good idea? Yes, this is fairly obvious. You could cast this as an application of a general principle that bombs blowing up in crowded areas is bad: you don’t need to stop and think or do any utilitarian calculus to jump over and get it done.
Should we have a prior that attempts to improve the extremely long term future make the world a better place?
The track record of these efforts are mixed. One can point positively to developments such as civil rights and the end of slavery, which so far have had long term positive effects on humanity. However one can also point to examples of utopian thinking that lead to unimaginable suffering, such as Stalinism.
However even these efforts are comparatively short term if we are talking about thousands or millions of years into the future. While things like religions have influenced humanity over thousand of years, the resulting societies can hardly be thought of as a stable or predictable consequences of the initial founding. Anthony gives a long list of example of backfireing effects in this article.
You’ve gotta be careful here that your reasoning is not circular here. For example, if you are an optimist, and your principle is that you believe that the positive influences of actions will beat the negative in the long term, you need to explain why you believe that. You cannot say that the reason is “I thought about the general effect of lots of different actions, and they seemed good”. This is just a restatement of our original problem, so it gets us nowhere.
I personally do not see a compelling way to argue that we should a-priori expect an action to yield good outcomes over a very long period of time, without reference to actual simulations of how that outcomes would affect the future. And I don’t trust those simulations at all.
Suppose Doug happened to read a trustworthy article saying “interestingly, the proportion of red fish to green fish in all the ocean happens to be quite similar to the proportion of red fish to green fish sold in pet fish shops.” But unfortunately the rest of the article was missing.
The question of the proportion of red fish in the ocean is very hard. But thanks to the intervention of the documentary, Doug has discovered that solving the hard looking problem is equivalent to solving an easy looking problem. As we already established, he can be confident of the university question, which will extend to the entire ocean as well.
For a shortermist equivalent, imagine that you are trying to help people in an impoverished neighbouring nation called Povertania, but don’t have a lot of good data or information on the effectiveness of individual interventions there. However, there is a very close upcoming election coming up in your own country, where one of the leading candidates, captain Genocide, has a campaign pledge to “invade Povertania, genocide their population, and loot their natural resources”.
You may be clueless how to help Povertanians in general circumstances, but you can be very confident that electing captain genocide will make their life substantially worse. If you find a cost effective way to prevent captain genocide from winning the election, then achieving this very short-term goal can effectively substitute in for the actual goal of improving peoples lives more generally. This works because we have a very strong reason to believe that the shorter-term crux point will have a serious effect on the longer term target.
When it comes to longtermism, this is a common defence. Sure, predicting and affecting the future a million years in advance seems impossible. But you can predict what humans will be like in the case of, say, human extinction: namely that that we will all be dead.
This is generally why longtermists tend to end up as anti-extinctionists. If your metric of success is “number of living humans over the next million years”, then it’s clear that preventing an apocalypse in the next 20 years will result in a better outcome than failing to do so. You might be totally clueless about how many humans exist in a thousand years if we prevent the apocalypse, but you know that the answer will end up being “0” if we don’t.
However, generally long-termists do not value “number of humans”, but some measure of the total flourishing and happiness of those humans. To make a concrete claim about this, we now end up in the position of proving that if humanity survives extinction, that it will result in good outcomes for humanity. This challenge is just as difficult as all the other ones we have already discussed.
It’s entirely possible that after preventing human extinction, the world ends up in a fate that is worse than human extinction. The application of this argument relies on an optimistic belief that if humans are alive, things will be happy and good. I do not believe we have the predictive power to prove this position. Skepticism becomes extra plausible when we think about the potential fate of non-human animals and of potential sentient AI systems, which could easily get significantly worse in proportion to human increases.
The other variant of this argument talks about “lock-in”, the concept that certain events will result in a future that is completely unchanging, generally due to some mythical omnipotent machine god that is totally due any second now. These scenarios are so speculative and frankly absurd that they fall beyond our veil of cluelessness discussed in the rest of this article, and so are of no help in my eyes.
Coming out of all of this, it seems to me like there are a lot of strong defences for believing in short-termist interventions, like funding malaria vaccines for people living in extreme poverty, when your scope of evaluation is also in the short term.
A negative outcome for a malaria intervention is certainly within the realm of possibility, but we have strong evidence from controlled trials that these interventions work, we have examples from the past of these interventions working, we have a clear scientific understanding of how these interventions work, and there has been a significant amount of in-depth research into all aspects of these interventions, including the impact of unintended sideffects. With these tools, we can chart the sea of uncertainty ahead of us, and confidently say that malaria vaccines are likely to be helpful in fulfilling short term goals of improving human lives.
In contrast, when it comes to the long term good, we face a challenge of prediction that is unimaginably more difficult, and we face it with tools that are greatly inferior. We have no scientific principles, prior experience, track records, or magical crystal balls that allow us confidence in our predictions.
This extends to the long term effects of short-term effective actions. It’s quite possible that saving the lives of impoverished villages will, through some quirk of fate, end up making the world of a thousand years hence vastly worth. Or it could go the other way and make the future vastly better through some other string of fate.
So if we are clueless about whether funding malaria vaccines will be good for humanity over the whole sum of time, then why should we do it?
Well, because we know funding malaria vaccines will be good for humanity over the short term. If we are clueless about the long scope, it shouldn’t affect our decisionmaking. So we should look at the shorter scope instead. It is fairly common-sense morality that saving a drowning kid is good, even if you didn’t do a full-scale calculation as to it’s ramifications over the next million years. Similarly, since we are not clueless about whether saving people from malaria is good over a scope of several years, we should do that.
In this article, I have constructed an analogy comparing the the cluelessness of a man reasoning about fish having never seen any to the cluelessness of long-termist altruists reasoning about an ocean of future possibilities without ever seeing them.
I have gone over four defences against this accusation of cluelessness, and shown how they can be used by short-term altruists to defend against cluelessness accusations, but are comparatively much weaker defences for long-termist reasoners.
From this, I make the argument that an altruist should in general prioritise short-termist causes over long-termist causes, due to the much higher chance of escaping cluelessness and actually knowing that you will make a difference.
I do not claim that I have covered every possible defence of long-termist beliefs here. I’m sure that other defences do exist, but I expect them to fall into the same pattern as the ones covered here.
The problem in this article is an separate one to my other post, on how the optimizers curse might guarantee the overestimation of speculative threats in EA. Both problems stem from a similar cause: the lack of modelling of estimator fallibility when decisionmaking.
I have also not gone into depth here as to the substance of my skepticism about EA’s ability to predict the long term future.That is another topic, and I have covered pretty extensively my disagreements with orthodox EA belief in countless articles on my blog. I think that forecasting the long term future is a ridiculously difficult task, and I do not believe that the effective altruist community is up to it. Nor, for that matter, is anybody else.
But I do believe that there are things in the near future that we can predict, and can change, if only we are smart enough to withdraw our gaze from the vast uncertain future, and focus on the people we can save here and now.
And do you think you should not be clueless about whether the near-term effects of your actions dominate (ex post) the long-term ones?
If you think you should, neartermism still does not follow without assuming something like bracketing (since it'd be indeterminate whether ex-post neartermism is true). So then, what you must do is defend some version of bracketing (against its problems), or some alternative.
If you think you shouldn't and believe the near-term effects dominate, it isn't clear how the unawareness argument against your position does not bite just as hard as the unawareness argument against longtermism. Why should we trust your best guess that the near-term effects of our actions dominate (ex post)?
(While I believe DiGiovanni in fact happens to believe the long-term effects generally dominate ex post, I think he is aware of the above, i.e., that indeterminacy on this question is enough for his argument to go through, and that this is why his argument actually does not assume longtermism, anywhere.)