This is a timed blog post. Every 5 minutes while writing this post, I had to stop to do 10 push-ups, and when I could no longer complete my required number, I had to upload it. My friend and I thought this would be a fun challenge, but it means that it will likely be less polished than some of my other output.
A few weeks ago, I wanted to improve at doing back-of-the-envelope calculations (BOTECs). So, I did a practice BOTEC and asked for some feedback. This post compiles what I learned: how to complete a BOTEC, what BOTECs are good for, when they’re valuable, and what their limits are.
Step 1: Identify the project
My BOTEC-ing journey begins with my time in the Generator Residency. At the beginning of the program, residents are asked to choose a project that they suspect will be impactful and execute on it for the remainder of the program. I had heard that BOTECs were a valuable tool for impact estimation, and I was eager to give them a try. So, I decided to analyze the idea of running joint hiring rounds for AI safety orgs. The idea looked something like this:
Aside from the main theory of change, there were other concerns about whether or not this project would be good for the ecosystem.[1] However, for reasons that will become apparent later, I don’t think this matters much. Besides, for the simplest version of our BOTEC, we won’t need to worry about that.[2]
Step 2: Identify the main source of impact
It’s good to have a clear theory of change (ToC) for any project you’re pursuing. In other words, you should have a clear understanding of what makes your project valuable for your end goals.
When you complete a BOTEC, the first thing you should do is identify the primary source of your project’s impact. Then, you should find an easy way to turn this into a unit that you can compare across actions. This is going to be the metric that you use to compare the value of projects. There are a few rules that can help you think about this:
The unit that I chose for my BOTEC was something I called “Gmult”, which stood for “Generator-skill-adjusted generalist time multiplier.” This could be interpreted as follows: for every 1 hour of Generator resident time put into this project, it saves Gmult hours for other generalists across the AI safety space (adjusted to the skill level of the median Generator resident). I chose this because it seemed like an intuitive metric for efficiency gains, and I thought it would show if I were obviously wasting my time.
When you’re choosing a unit, you should think carefully about the goal of your BOTEC. Are you trying to compare two specific projects? Are you trying to communicate a true thing to a wide audience? Are you trying to understand a topic better? Your answer to these questions might change what unit works best for you.
Step 3: Estimate the shape of your formula.
The next thing you need to do is figure out which components make up your final answer, and determine how they relate to each other. The units of your components are very important, so you should specify them clearly if you intend to communicate your results with others. Here are some things to think about:
Additionally, you should choose units that help you quickly get to your result and are easily measurable. For my BOTEC, I chose a few simple parameters:
My resulting formula was:
Gmult = n * Smult * %TS
You can double-check this formula by plugging in different numbers and checking the output. For example, if you save 100% of the hiring time at 3 orgs, and Smult is 1 (the assumption here is that Generator residents are just as skilled as workers at other orgs), then joint hiring rounds produce 3 skill-adjusted generalist hours for every hour put into them.
If this answer seems obviously wrong or confusing, you might want to double-check your formula. For example, maybe Generator residents would be slower hiring managers than other professionals. You can add another parameter could be added to adjust for this.
Step 4: Determine a plausible range for your numbers
Next, you need to figure out what numbers you should plug into your formula. This is where everything can fall apart if you’re not careful. The output of your formula is only as accurate as its inputs, so this is important to estimate these correctly. Here are the ranges I used for this project:
I don’t think these numbers were especially accurate, but I don’t think they were obviously wrong. Had I wanted to come up with a more serious estimate, I would have spent more time carefully thinking about these numbers. However, my goal was only to learn BOTEC-ing, so I didn’t invest much time into this.
Step 5: Estimate and interpret
This led to a Gmult range spanning from 2.16x to 16.2x. My initial intuitions about these numbers went something like this:
These estimations led me to believe that this project would not be worth pursuing on impact grounds alone.[3] This evaluation also matches the story of one of the other residents: they were initially interested in completing this project, but they pivoted to working on sourcing candidates for top orgs directly. This turned out to be a much more efficient use of their time and one of my favorite Generator projects.[4]
I got several things out of completing this BOTEC:
All of these qualities seemed pretty good and made me feel better about the project-selection process.
However, despite their benefits, there are several reasons you might want to distrust BOTECs:
I think the points made here are important enough that if you ever want to think about doing BOTECs, you should probably slow down, reread this, think carefully about each point, and maybe reconstruct them all from memory. It’s tempting to skim list-posts, but it is very important to understand the limits of any epistemic tool you use.
So, given these weaknesses, are BOTECs worth doing?
If you’re choosing between projects to work on, most of the value of BOTECs comes from gaining a clear conception of where your project’s impact comes from. So, I would recommend thinking carefully about your theory of change, and if you’re confused about that, you should determine what your unit of impact probably is.
In the early days of Generator, one thing I noticed was that many people didn’t spend enough time thinking carefully about where the primary impact of their projects would come from. This ultimately cost them lots of time and emotional strain, as they only realized a project wouldn’t be very good after several more hours of work. I suspect that some of them could have avoided this trap if they thought more carefully about the main unit of their impact.
Aside from that, I think you should only worry about completing a full BOTEC on projects you plan to spend more than ~24 hours on. As a general rule, BOTECs take longer than you expect to get right. You will probably learn much more and do more good things by just trying to go do things.[5]
More broadly, things are only as valuable as how much they help you achieve your goals. It’s easy to get distracted by the allure of BOTECs: math is fun, and it can feel rewarding to clarify your thinking on every small issue you encounter. But don’t confuse this for impact.
Maybe it makes hiring choices too strongly correlated, and maybe orgs already do a good job of recommending people to new roles
If it ultimately mattered for the end result, I would have found a way to adjust for it.
It might have been worth it for upskilling reasons, but I’m kind of skeptical of this.