I put together some thoughts about cluelessness for the Cluelessness Critiques competition. The essay wasn't accepted to be part of the main competition, since it doesn't engage closely with Anthony's argument, but I figured I'd post it anyways. In retrospect, what I wrote might be a bit dense/elusive, but I hope it can still be interesting.
My constructive response to the problem of Cluelessness: suppose we accept Anthony's conclusion — that we have no impartial altruistic reason to prefer one action to another based on the good-ness of its consequences. An impartial altruist still might be able to compare how reversible the consequences of different actions are, and to prefer actions which are more reversible, since they might lead to better outcomes in the long-term.
Anthony DiGiovanni has argued that impartial altruists cannot justifiably prefer one action to another — since we are clueless about their long-run consequences. Nonetheless, we are tasked with deciding how to do good in the world. So how should an impartial altruist choose how to act?
They could choose to do nothing, but that has its own consequences. One decision-making strategy is Bracketing, basing decisions on the consequences we are not clueless about. This may lead us to favor nearterm interventions. On the other hand, the problem of cluelessness could lead us to favor Longtermism, as Hilary Greaves has suggested. The argument from cluelessness to Longtermism is roughly as follows: long-term effects supply most of the expected value of our actions; but we are clueless about the long-term effects of near-term interventions; so we should prefer interventions which are aimed at the far future, since their long term consequences are more predictable (for example mitigating existential risk: biosecurity, nuclear, and AI risk).
I will argue that impartial altruists should favor reversibility: preferring interventions which are more steerable in the long term. If we are indeed clueless, then some of our actions will have significant negative consequences. To prepare for this eventuality, impartial altruists should prioritize actions which can be more easily reversed in the event that unanticipated negative consequences become evident.
To reverse an action is to undo all its consequences. An action is reversible if it can be reversed in finite time. An action is irreversible if it cannot be. However, given that the flap of a butterfly’s wings in Brazil can set off a tornado in Texas, it seems like every action is irreversible in practice, since we can expect the cascading consequences of any action to extend untraceably far in all directions. If we can’t keep track of all the consequences, or if they are effectively infinite in scope, we’re unlikely to be able to reverse them. Nonetheless, we intuitively sense that some actions are more reversible than others.
Consider a glass resting at the edge of a table. If we hit it on the left it will fall off the table and shatter. If we hit it on the right, it will fall on its side and remain unbroken. Hitting the glass on the right is more reversible than hitting the glass on the left in the sense that it’s easier to pick up the glass and place it back upright than to collect all the broken pieces and glue them together. For many actions, subsequent actions can change their consequences. For both a glass lying on its side and one that’s shattered on the floor, it’s not possible to fill it with water from a pitcher. But if the glass is restored — either by gluing or placing upright — then it can be filled. That is to say, many actions are partially reversible, and unequally so.
The glass example is interesting partly because one action can restore most of the functionality of the glass — gluing or placing upright. There are other actions which cannot be reversed in this wholesale fashion. For example, consider spreading gossip. Jane tells Jesse that Jack is going to quit his job. In fact, Jack is not going to quit his job, and Jane is making it up. Nonetheless Jesse tells five other people. To undo her original act, Jane will need to speak to Jesse as well as each of the five people he spoke to individually. Considerably more than the single conversation she started with. (Although it’s possible that undoing the act could also follow the same propagation logic, so that speaking to Jesse would be enough since he would follow up with the people he spoke to).
There is often an asymmetry between the original act and the undoing action — smashing the glass is easy, repairing it is hard. Spreading gossip is easy, taking it back is more difficult. There are also actions with the opposite asymmetry. Say we erect a statue. That requires, at a minimum, sourcing materials and engaging an artist. Undoing the act may be simpler — for example smashing the statue. Or, if it’s an ice sculpture, the action undoes itself: the statue melts away.
In the next few sections I’ll attempt to be precise about what reversibility is, what its problems are, and why nonetheless impartial altruists might choose to make decisions based on it. Unfortunately, being precise in the way I’ve chosen to be, there’s a real risk of philosophical nonsense. So take the argument with a grain of salt, I’m mostly pursuing it as a means of dredging up problems with the idea of reversibility. The core intuition — that impartial altruists should expect some of the actions they take to have negative consequences and should prepare accordingly — should survive despite any insufficiencies in the argument.
Is reversibility a property of an action’s consequences? I defined (true) reversibility as reversing all of an action's consequences; however this may not be consistent with how we actually think about reversibility intuitively.
When, for example, you’ve made a mistake — say knocking a glass off the table — and want to reverse that action, you do not actually sit down and enumerate all the consequences of the glass having been broken (you can’t drink water out of it, you can’t drink milk out of it, you can’t drink orange juice out of it, etc.) and address them each individually one by one. It’s more likely you will think about restoring the original circumstances — when there was a functional glass on the table and no broken glass on the floor.
When it comes to reversing an action, we often don’t actually think so much in terms of its consequences and more in terms of its particulars. In the glass example, these are the glass, the table, the person who hits the glass, etc. — all the things involved in the action. When the glass is shattered and we want to reverse the action, the objective may be to make sure that the person has the same number of glasses they had originally. That has more to do with the particulars involved in the action than it does with abstract concerns about the state of possible worlds.
I’m automatically a bit suspicious of any measure which relies on imagining possible worlds, like calculating long-term EV does, since doing so relies a bit too much on the human imagination, which is deeply psychological and — dare I say — not impartial. That is the key advantage, to my mind, of thinking of reversibility in terms of particulars rather than consequences. In particular, we may be less clueless about reversibility in the particulars sense compared to reversibility in the consequences sense, since it has more to do with concrete objects, and less to do with imagining possible consequences.
Then, (partial) reversibility is the cost of restoring an action's particulars to something near to their original circumstances.
It may be easier to calculate the cost of restoring particulars — which are concrete, compared to consequences which are abstract. Moreover, they can be reasoned about empirically. For example, say that you erect a statue. If you were to reverse that action, you could experiment with different methods of destroying it — melting it down, throwing it in the ocean etc.
However, the notion of partial reversibility is narrower than the definition of true reversibility. I’ll discuss how the two definitions interact in the next section.
How do the particulars and consequences definitions of reversibility interact? And, moreover, how are an action’s particulars and consequences related? One way of seeing how they interact is by thinking about an action’s scope.
The scope of an action consists of the particulars which are directly involved in the doing of the action, or can be suspected to be affected by it. In the glass example, this is the glass, the table, and the person knocking over the glass etc. In the gossip example, Jesse and Jane are in the scope of the original action, but the five people who Jesse tells are not.
The direct consequences of an action are the consequences which are in the action’s scope. That is to say, consequences which concern the particulars in the actions scope. For example, the intended effects of many actions are direct consequences.
One question we can ask ourselves is whether it’s enough to know that you can reverse an action’s direct consequences to say that the action is reversible in general. This is roughly (but not perfectly) equivalent to asking whether it is possible to infer whether an action is truly reversible in general from its partial reversibility.
One problem is that some actions exceed their scope. When Jane gossips to Jesse and he tells other people, the direct consequences are those which have to do with Jane and Jesse. Namely that Jesse believes that Jack will quit his job. However, the fact that five other people now believe this as well is not a direct consequence. So, when Jane tries to reverse the action, it may not be enough to reverse the direct consequences, since speaking with Jesse alone may not be enough to resolve the situation. The action has exceeded its scope, since reversing the action requires particulars which were not within the scope of the original action.
On the other hand, it’s possible that second-order consequences will sometimes depend on direct or intended consequences. For example, a consequence of distributing malaria nets is to save lives which would otherwise be taken by malaria. A second order consequence — a consequence of saving lives — is that distributing malaria nets may cause population growth. Since the second order consequences depend on the intended consequences, if bed nets stopped being distributed, it’s possible that the lives saved would stop increasing with the corresponding effect on population size.
It seems reasonable to suspect that an action whose direct consequences are thought to be reversible is more reversible in general than an action for which that’s not the case. To bring the same suspicion to the two definitions of reversibility: an action’s partial reversibility may be correlated with its true reversibility. Although we can’t know for sure, an action which is partially reversible may be more likely to be truly reversible as well.
It may be useful to quantify reversibility — the cost of reversing an action — and to relate it to expected value (EV, an action’s goodness).
We have been talking about reversibility as the cost of restoring an action’s original circumstances. If a glass is shattered on the ground there is a cost in time and resources associated with putting it back together. It may also be wise to think about costs besides financial cost.
For example, perhaps we are, actually, fairly uncertain about whether the original circumstances can be restored. In that case, maybe there should be a penalty for that uncertainty — or that reversibility should be given as a range of costs rather than a single number, similar to how EV is sometimes measured as an interval.
There may also be a moral cost associated with reversing an action. Unfortunately, one consequence of being an impartial altruist is that the actions you are typically interested in are good — very good. This means that reversing them can seem a little bit evil. Imagine the Against Malaria Foundation stockpiling mosquitos infected with malaria to release in case it turns out the whole ending malaria thing is net negative in the long run. Obviously, this is not the sort of thing I’m recommending. There may be a real moral, reputational, or existential cost to reversing an action, or preparing to reverse an action, which might need to factor into calculating the cost of reversibility.
Also, an action is not good because it is reversible. Knowing that an action is reversible does not mean it’s also good — reversibility isn’t sufficient for the creation of moral value. In other words, reversibility is not an engine of goodness, that comes from somewhere else. Reversibility may factor into our decision making, but mostly as a way of deciding, on the margin, between two actions which are already thought to be good.
For example, take the expected value of an action α, say EV(α) (calculated using bracketing or some other approach), and subtract from it the cost of reversing the action R(α), perhaps with an additional parameter λ, to control the relative importance of each of these terms in deciding the outcome.
EV(α) – λ R(α)
Under this criteria, actions which are highly irreversible — for which the cost of reversal is astronomically high — will be deprioritized.
Reversibility can sometimes seem a bit evil, or potentially increase existential risk, or be something we are a bit clueless about. We could also worry that if an action is reversible, maybe it wasn’t all that meaningful in the first place — the only truly ‘effective’ interventions are the ones that are highly irreversible. However, I’m a bit suspicious about this. Consider measles which was declared eliminated in the US in 2000, and robustly good. But maybe not irreversible, considering recent outbreaks.
On at least one point, the reversibility argument seems at odds with Anthony’s Cluelessness sequence. Namely, reversibility may as well be a heuristic for making decisions, and Anthony has argued that heuristic judgements about ‘goodness’ aren’t enough. However, Anthony mostly means this for intuitive judgements about goodness. Since reversibility has implications for an action’s consequences, namely bad consequences may not be permanent, then we can truly c-prefer one action over another on grounds of reversibility.
Reversibility also seems to rely on the existence of beneficent future people who will reverse actions when they turn sour. Can we trust future people to make choices which are good for the long term future of humanity? I’m not sure. Although it may actually be a good thing, in and of itself, to empower future people to do good. Because of reversibility, they can either choose to live with the altruistic interventions we’ve done, or to reverse them.
For example, in 1998 the economist Amartya Sen won the Nobel prize, partly for developing the capabilities approach, which argues that human well-being should be characterized by what people can actually do and be, rather than utility, money or other measures. By this measure, reversibility, which empowers future people by giving them more control over the consequences of the actions we initiate today, may be good for that reason alone. Although the benefits of this empowerment may not be evenly distributed.
Reversibility is a bit problematic. Then again, so is EV, as Anthony’s sequence on cluelessness will attest.
Anthony’s conclusion says that we have no impartial altruist reason to prefer one action over another. That, on its own, wouldn't be so much of a problem. We can imagine impartial altruists maintaining a diverse portfolio of philanthropic ventures expecting most of them to be duds and some to be great successes, like venture capitalists. The problem is that some philanthropic interventions may fail catastrophically, leading to massive negative impact which could potentially outweigh the good done elsewhere, so the venture capital analogy might not be appropriate.
For example, one reason Habryka gave for closing the Lightcone offices in 2023 was a growing suspicion that the ecosystem he'd been building infrastructure for was net harmful — not least because EA and rationalist ideas helped found DeepMind, OpenAI, and Anthropic, plausibly shortening timelines to AI catastrophe. We should expect that, in the future as well, some initiatives impartial altruists are involved with will have net negative effects.
That’s not to say that the project of doing good is hopeless or should be abandoned, just that it should be approached with humility. Impartial altruists should operate under the assumption that some of the interventions they plan — no matter how well justified — may have negative net impact. We might call this the fallibility of impartial altruists.
How does reversibility compare to other responses to the fallibility of impartial altruists? It may be more robust than ‘better estimation of EV’ since reversibility is not about making sure we’re wrong less of the time, which is difficult to optimize enough to eliminate fallibility, but making sure we’re better equipped to respond when we are wrong. It also may be a better fit for the EA culture compared to other similar responses like ‘risk aversion’.
Moreover, an impartial altruist may justifiably prefer actions which are more reversible. Reversibility has implications for an action’s consequences. An action which is highly reversible is more likely to have any unanticipated future negative consequences reversed in the event that they become manifest. So an impartial altruist, who makes decisions based only on an action’s consequences, may justifiably prefer actions which are more reversible.
In short: reversibility is an approach to managing the inevitable negative consequences of well-intentioned good actions. It’s not perfect, but is competitive with other similar strategies. Perhaps at this point the best next step would be to see what sorts of interventions the reversibility argument might actually recommend.
It pains me to say it, but knowledge creation is one of the most irreversible actions there is. Once the cat’s out of the bag, you can’t put it back. It’s like gossip — inevitably new knowledge, if it's interesting, will spread and have potentially heavy long-term consequences. New ideas or research may have unintended consequences. Think of things like technical AI safety, which could contribute indirectly to advances in model capabilities and so increase x-risk. Or researching alternative proteins. On the other hand, investing in new and better knowledge might be one of the most reliable ways of empowering future actors to do good better, potentially for a very long time — if that knowledge persists and continues to be useful for many generations. Unless you don’t believe that humans actually make better decisions when they have more knowledge.
Lock-in, where a single ideology, culture, or artificial intelligence system takes control of the world, seems fairly irreversible.
Letting a species go extinct is a highly irreversible action, especially if no genetic material from the animal is saved. You could say the same thing about allowing information to be lost — the early internet for example. Language death.
Disease eradication, which may seem highly irreversible, is actually highly reversible. Diseases thought to be eradicated often resurge — I already used measles as an example, which was declared eliminated in the US in 2000, but today there are more cases. Stockpiling a small amount of the genetic material of a disease might prevent its eradication from being irreversible.
Or disease intervention which stops short of eradication. Recurring inputs like bednets, deworming, micronutrient supplements are reversible (with perhaps a high moral cost).
Cage-free reforms are highly reversible — the legislation can be repealed.
Stockpiling (PPE, vaccines, food reserves, seed banks) is reversible. So is archiving — like what the Wayback Machine on Internet Archive does.
In general, building things which can either be destroyed, abandoned, or deprioritized is fairly reversible. For example building bunkers, refuges, or other civilizational insurance.
Even the most careful and well-intentioned impartial altruists will be wrong — sometimes disastrously wrong — some of the time. That shouldn’t cause us to cease being careful, rigorous, and well-intentioned, but is reason for some humility. Reversibility is a security measure to prepare for the cases when impartial altruists are wrong. By choosing the actions which are easier to reverse in the first place, the risk of irreversible negative consequences is slightly mitigated.
It’s still a bit disheartening to think about impartial altruists carefully planning actions which do good extremely effectively, only for them to one day turn sour.
I see some comfort in the final scene of Goethe’s Faust. Arguably the most famous piece of German literature, it ends with Faust ascending into heaven as a spiritual chorus delivers a final philosophical message:
“Alles Vergängliche ist nur ein Gleichnis.”
All that is transitory is but a symbol/metaphor.
It’s a powerful sentiment. One which inspired an entire essay devoted to the scene by Adorno (“Zur Schlußszene des Faust” or “On the final scene of Faust”). In the 19th century, the phrase was used as lyrics in music by Schumann and Liszt. In the early 20th century, the text was set by Mahler in his 8th symphony — nicknamed the ‘symphony of a thousand’ for the number of performers required to stage it at full scale. It’s one of the largest-scale works in the repertoire. Alles Vergängliche is the final, climactic chorus. Mahler drops the orchestra and the line is sung a capella.
It’s hard to imagine a more powerful endorsement of a phrase that is, admittedly, a bit opaque. In my reading, it reflects on how transient things may refer symbolically to a larger project. The original meaning was obviously a bit religious, but I think it survives outside that context. For example, I am thinking of the Kantian J. C. A. Grohmann who said ‘the history of philosophy is the end of all philosophizing.’ Or, more bleakly, in the words of the 17th century writer René Rapin, ‘We have seen these philosophies being born, and we shall see them die.’[1] Individual philosophizing is typically proved wrong or eventually overshadowed. Nonetheless it pushes forward the larger project of philosophy.
We might say the same thing about altruism — the end of doing good is the history of doing good. Although impartial altruists may be fallible, the good we do will inform future people trying to do good. Hopefully what we do today, in addition to making the world a better place, will inspire and empower future people to do good as well — to the benefit of all.
Both as cited in The Descent of Ideas by Donald R. Kelly