Disclosure: I work at FutureSearch, which makes the tool this post is about. I'm sharing it here because I genuinely think decision forecasts are useful for effective giving. Our results are freely available, and our app is free to try.
Donors frequently face tough decisions about how to best distribute their limited funds. Whether deliberate or implicit, they're essentially making a series of forecasts about the effects of their donations. For instance, a donor might consider, "If I donate $X to organization Y, what's the likelihood they'll achieve outcome Z?"
However, naively forecasting a question like this brings in factors I don’t want. For instance, my donation itself is evidence that the project will succeed because I wouldn't have funded it if I thought it would fail. Instead I care about the causal effects of my donation, not the correlational ones. That's why we developed a technique we call decision forecasting, which compares the likelihood of achieving the outcome with and without my grant.
To demonstrate how it works, we ran our AI forecaster on 18 open grant proposals at three funding levels each to help inform which ones are the most funding-constrained.
When forecasting, choosing the right outcome is often as important as the forecast itself. As a donor, I care less about whether the organization meets its stated deliverable and more about the impact of my donation. However, impact is difficult to measure, and different people can reasonably disagree about the impact of an outcome.
A helpful proxy is predicting uptake by influential members of the community. While not it's not the same as impact, uptake is a strong indication that the community finds the work valuable. Therefore, for this study, we considered the question:
By December 31, 2028, will a named outside actor take a costly action, documented in public records, that stakes something on this project's work produced after August 1, 2026?
This framing is generic enough to apply to the large majority of AI safety proposals, allowing us to compare them on a shared scale. But for the forecasts to be concrete, we need to specify what actors and actions count. For ControlAI, the actor is a sitting parliamentarian signing a public campaign statement. For Transluce, it’s a frontier lab or a government AI safety institute documenting the use of Transluce's tools in an official model evaluation. Passing citations and social media engagement don't qualify.
However, this framing has its limitations. It doesn't cleanly capture community resources like Mox or projects aimed at raising awareness like Doom Debates. For this exercise, we chose to exclude proposals which don't suit our chosen outcome. Of course, you can choose your own outcome to forecast.
We compiled 18 public AI safety funding proposals from Manifund, grantmaking.ai and other sources which were open as of July 31, 2026.
For each one we considered three donation amounts: nothing, a partial amount, and the amount to meet their funding goal (as of July 31). Importantly, these are the amounts you would hypothetically donate. If you choose not to give, other donors may step in. The $0 forecast serves as an unconditional baseline of the outcome's likelihood. And comparing the forecasts for different funding levels indicates the causal effects of your decision.
This table shows the proposals we considered and their forecasts. The Delta column is the percentage-point gain from $0 to the full funding ask.
Note: I sorted by delta because it's the column that surprised me most. The order does not imply worthiness since outcomes differ in impact and the funding amounts differ in size.
| Organization | Outcome | Ask (partial, full) | P at $0 | P at partial | P at full | Delta |
|---|---|---|---|---|---|---|
| ControlAI | 250 non-US G7 parliamentarians listed on its campaign statements | $250k, $1M | 39% | 54% | 70% | +31 |
| Evitable | 2 national politicians or civil-society orgs join or cite its campaigns | $300k, $1.49M | 57% | 71% | 83% | +26 |
| Apart Research | 2 unaffiliated safety actors build on or cite its new research | $200k, $820k | 37% | 47% | 59% | +22 |
| Token taxes | Its policy memo cited in an official government document or proceeding | $100k, $370k | 41% | 52% | 62% | +21 |
| Palisade Research | New findings cited in 2 congressional hearings or US government publications | $250k, $1.13M | 40% | 48% | 59% | +19 |
| AI Digest | 3 citations by government bodies, national outlets, or policy processes | $150k, $682k | 47% | 55% | 65% | +18 |
| Center on Long-Term Risk | 2 outside publications substantively build on its new research | $100k, $400k | 38% | 45% | 55% | +17 |
| GPAI Policy Lab | An external actor uses, pilots, or cites its proof-of-training outputs | $250k, $2.47M | 11% | 16% | 27% | +16 |
| Foresight Institute AI Nodes | 2 residency outputs adopted by unaffiliated safety actors | $125k, $503k | 16% | 21% | 29% | +13 |
| Standalone world-models (Thane Ruthenis) | New agenda outputs published and adopted by a recognized safety actor | $50k, $200k | 21% | 29% | 34% | +13 |
| Understanding Trust (Abram Demski) | 2 safety actors adopt his tiling-agents research | $50k, $144k | 22% | 28% | 33% | +11 |
| Transluce | 2 frontier labs or safety institutes document its tools in official evals | $500k, $1.96M | 21% | 25% | 31% | +10 |
| AI Safety Camp, 12th edition | 2 camp outputs adopted by recognized safety actors | $20k, $50k | 16% | 20% | 26% | +10 |
| Forethought | 2 official policy documents cite its new research as a basis for action | $500k, $2.9M | 12% | 15% | 21% | +9 |
| AI Futures Project | New materials used in 2 official government settings | $150k, $456k | 43% | 47% | 52% | +9 |
| PauseAI US | 10 sitting members of Congress sign its public letter | $100k, $450k | 7% | 10% | 15% | +8 |
| Timaeus | 2 unaffiliated safety actors use or extend its methods | $150k, $589k | 46% | 47% | 49% | +3 |
| Tarbell Center for AI Journalism | Supported journalism cited in 2 official government proceedings | $50k, $171k | 78% | 79% | 81% | +3 |
Click here to view all the forecasts and their detailed rationales. Our AI forecaster ran an average of 78 searches and read an average of 39 unique pages per proposal.
These are probability estimates from automated web research, produced without contacting the applicants. They concern one particular outcome per proposal and are not assessments of the applicants nor verdicts on their work.
The forecasts focus narrowly on whether the outcome will be met. As a result, red flags which barely move that probability get less research attention than they deserve in a funding decision. FutureSearch can help discover and prioritize funding opportunities, but we recommend further diligence before writing any checks.
Rationales can contain errors. Please let us know if you spot any mistakes.
We welcome input from all donors, both large and small! Even though we focused on AI safety proposals and donations with six to seven figures, the same approach works for other cause areas and amounts.
You can start with our results and ask follow-up questions. Or start from scratch with your own proposals, outcomes and donation amounts.