The AI safety movement should push itself to be dramatically more transparent to the public.
To date, the AI safety movement has been one of the strongest forces for clarity and wisdom in the world. The movement has been prescient on the subject of concerns from existential risk, seriously grappling with outcomes others dismissed as sci-fi nonsense.
Society is now waking up to the potential threats of advanced AI. I understand that many in the movement are feeling the crunch, and thinking more carefully about optics and what they publish. Even so, acting transparently is more important than ever.
Why transparency?
If you want labs to be transparent, you should be transparent too. AI 2040 proposes “Total Research Transparency” for labs to open up their research, algorithms, LLM weights. Safety should do likewise. Model good behavior, to convince labs that this is an acceptable and correct way to behave.
I think AI safety people are unusually virtuous; you should display that virtue. “Nor do they light a lamp and then put it under a bushel basket; it is set on a lampstand, where it gives light to all in the house.” (Matthew 5:15)
Transparency ties you to the mast, forces you to be virtuous. Famous maxim: “Act as though what you do might end up on the front page of the NYT”. And, what better way to enforce that than to publish everything you think and do?
You can’t keep things private anyways, given stylometry and cheap intelligence. Actions cast a shadow in the world, and AI will be able to detect that shadow, reconstruct that action. (More here.)
Transparency was a cornerstone of this movement. It is part of what drew me (and many others) to engage and buy into the beliefs of AI safety. Public writings, auditable spreadsheets. Sticking with transparency would demonstrate that the movement has integrity and self-consistency.
501c3 public charities are obliged to some transparency already, around financing and executive salaries. And, ~all AI safety orgs abide by the letter of the law. But perhaps, you should be exemplary with regards to its spirit.
It is a duty of powerful actors to be transparent, and AI safety is becoming powerful. Society already asks for transparency from our political leaders, our government, our labs and megacorps. AI safety wishes to influence major actors and pass sweeping regulation. “Dress for the job you want”; prove yourself worthy of this role.
Transparency helps with internal coordination. It scales well. AI safety is about to undergo hypergrowth. It’s no longer 200 people in the Bay Area who all go to each other’s parties. Knowing what other people think and are doing, regardless of who is in or out of particular group chats, will be key to scaling up.
Transparency helps with external recruitment and fundraising. To date, a lot of AI safety people have come to the movement via public writings and transparent reasoning. Many more great people will join as well, if you continue this way. (If transparency is good for recruiting, does that mean that only public comms-focused orgs like 80k and Bluedot should be transparent, versus our research and policy orgs? I’d argue no.)
The rise and fall of “Open Philanthropy”
Any accounting of transparency in AI safety must begin with the saga of Open Philanthropy.
Givewell began as a strong force for transparency among nonprofits. Holden and Elie published constant updates, maintained a listing of their own mistakes, made their spreadsheets available for public critique, and engaged in good faith with commenters. Take a look at the Givewell Blog circa 2007 for a taste of this; I find their attitude beautiful and inspirational. This transparency was key to the early EA movement, and I believe this helped to convince Dustin Moskovitz and Cari Tuna to start funding Givewell with major amounts of money. Together, they started Givewell Labs to explore causes beyond global health, which became the behemoth Open Philanthropy.
Sadly, this golden era did not last. In 2016, Holden published a major update on how they’re thinking about openness (mostly: less open, due to its costs). Beyond their stated reasons, I suspect that once Good Ventures (Dustin & Cari) became more committed to funding OpenPhil, OpenPhil just had less pressure to continue making its thinking visible to the broader public.
From there, things only became less transparent. Fewer public spreadsheets and debating in comment sections. I expect Holden remained a major driver for transparency; I greatly appreciate Cold Takes for this. But Holden left OpenPhil in 2024. And then in 2025, OpenPhil gave up on being “Open Philanthropy” altogether and renamed itself to Coefficient Giving.
After some reflection, I consider this sequence of events to be a failure of principled thinking. I hesitate to say this because Holden himself was (and remains) one of my personal heroes; he’s certainly in the top 5 of people who have shaped my views. (Others include Scott Alexander, Paul Graham, and Jesus). But: I think that Holden and others didn’t properly honor the role of transparency, in the growth of Givewell and later OpenPhil. Phrased extremely uncharitably: being transparent is what got CG to where it is; to discontinue it now is a bait-and-switch.
Does CG itself owe transparency to anyone beyond its funders? I tend to think yes. It owes it to the AI safety ecosystem: CG also draws from a scarce talent pool for its hires, and solicits lengthy application processes from its grantees. And obviously, CG operates as a public charity aiming to benefit the world. If you aim to serve the public, you should engage in dialogue with it.
(One could make a libertarian-ish argument that CG is free to only publish whatever it pleases. I’m sympathetic to this kind of reasoning; in that case, I wish that donors, talent, and grantees would vote with their feet. Hence this essay.)
Beyond internal transparency, I think CG should push the movement for transparency given its outsized role in the ecosystem. CG has been the biggest funder of AI safety, representing more than half of all dollars moved to date, and is also poised to grow rapidly given Dustin’s investments and others (eg Anthropic employees). And, while CG disclaims leadership of either AI safety or EA, I see this as a dereliction of duty; its funding has made these movements what they are. Grantees, and the broader movement, take cues from how CG itself behaves.
Other times when I’ve been disappointed by lack of transparency
- Many orgs not detailing where their funding comes from
- This has lately been coming back to bite EA, eg people casting aspersions on METR et al. I think these aspersions are mostly wrong — I believe METR has made large sacrifices to ensure its funding sources were independent as opposed to coming from CG — but I don’t expect my words to be persuasive to these critics. I think METR instead should just disclose how much each donor gave, and when.
- Coefficient Giving (and other funders like LTFF) for no longer publicizing grant writeups or rationales
- And, those are the better ones, for ever having published writeups at all! Other funders like Longview and Macroscopic are far more opaque, not even publishing the amounts to grantees.
- I hope that Lightcone Commons will aim for more transparency than this, though I don’t know what their current thinking is.
- Funders asking to remain anonymous (including, donors to Manifund)
- (some) AI safety leaders not writing much about what they think
- People hesitating to publicly associate with EA or AI safety
- Especially while still seeking their funding; especially in the lean times, post-FTX but pre-Anthropic dollars
- Constellation Slack, and the “internal Google Doc” meta
- Even worse: disappearing Signal group chats as a default form of communication
- Sorry for the controversial example, but: consider whether the same tactics used by FTX, are the ones you wish to employ now.
- Some political actions that the movement has considered or is currently doing
- Eg the AI safety volunteer campaign for Alex Bores. (Contact me for my writeup on this.)
- CG employees hesitant to say anything, needing to run things by their media team
- Similarly, Anthropic employees hesitant to say much in public (for fear of, leaking research secrets?)
- I appreciate that OpenAI seems somewhat more loose on this front
- Anthropic & OpenAI hiring a bunch of my favorite writers out of AI safety & EA, and then them going quiet. Eg Holden Karnofsky, Joe Carlsmith, Xander Balwit, Dean Ball, David Oks.
- This is mostly explained by these individuals becoming very busy, as opposed to anything sinister. And yet, I can’t help but see this as strip-mining the public commons.
- The Blog Revival Project is our (half-joking, half-serious) attempt to push against this.
Times where I have appreciated transparency
This is a shorter list than the above, but to be clear, I think AI safety does much better on transparency than almost any other comparable movement or community. I critique the movement because I have hope that they will improve.
- Richard Ngo’s recent history & retrospective on AI alignment
- MATS’s funder disclosures, among those of many other AI safety orgs
- Ryan Kidd commented, “You should also mention Alex Turner, Daniel Kokotajlo, Jacob Coxon, etc. As another example of radical transparency, GiveWell and Animal Charity Evaluators publish their board meeting minutes publicly.” (and linked to more transparency practices of MATS eg this, which I appreciate.)
- The widespread practice of listing the identities of employees on AI safety org websites, which I hope never goes away
- In background: anyone who writes anything on public internet forums (including LessWrong, EA Forum, Substack, X, or personal or organizational blogs, including Anthropic and OpenAI’s)
- AI safety leaders like Paul Christiano and Daniel Kokotajlo and Katja Grace, and more broadly AI leaders like Dario Amodei and Sam Altman, writing about their views on AI on blogs and conducting public interviews.
Manifund’s stance on transparency
To start, the basic premise of Manifund is transparent & fast grant applications. Anyone can post an application on the public internet; fund an application they like; comment about the merits or demerits of a particular proposal; see where and how money moves.
To practice transparency ourselves, we make our source code open; our data wholly available; our finances for anyone to inspect. Beyond our website and newsletter, an unusual amount of our thinking is available on our public Notion.
We’ve recently built tools to increase the transparency in the AI safety & broader EA movement. Trace is our attempt to improve funding transparency by listing every single grant made in the space. The AI safety funder bulletin is our attempt to explain what different funding orgs are up to.
We hope to go farther. We’ll soon be publishing the salaries and roles of all Manifund & Mox staff. We hope to publish more of our own thinking: on funding, AI safety, and what a flourishing future might look like. Movement-wide: we hope to publish analysis of specific orgs, funders, or grants, similar to Givewell of old.
Reasonable steps for transparency AI safety people should consider now
I think all of these are pretty unobjectionable:
- Disclose where your money came from, how much, and when
- Publish salaries, either anonymously in a levels.fyi style, or as openly as Jeff Kaufman
- Share grant applications in public — definitely after receiving funding, but maybe before, too
- This can also just be good for your org! We published our FTX Future Fund application, before it was approved. Ian Philips saw it and reached out, and became Manifold’s first hire, and later cofounder
- Share where your personal donations are going (especially to political candidates), and where your volunteer efforts are going
- Coordinate in public rather than in private, or at least disclose coordination after-the-fact (such as for the Pacing the Frontier letter and the Jacob Coxon resignation).
- Write more. Or, if writing is too hard, podcast more.
- Host regular “ask me anything” sessions
- Speak frankly about your intellectual and emotional influences
- Host “mistakes” pages like ACX and Givewell; post retrospectives after major issues
- Disproportionately fund people and orgs who are being transparent in public, rather than just those offering information to you-as-funder
Extreme radical transparency (which, might still be good)
An ideal version of Manifund might do these things; we’re not there yet, unfortunately.
- Opening up the full log of all payments into and out of your org
- Opening up all your internal communications to the public: emails, slack, google docs
- Sharing the full log of every LLM conversation and agentic trace used by your org
- (More broadly, I suspect it might be good if some regulation forced every lab everywhere to publish the logs of every interaction with an LLM, both in training and in deployment. At least, it would solve the problem of “eval awareness”.)
- Having every employee wear bodycams while “on the clock”, and/or broadcasting their screen via Twitch
- Note that our society already asks police and congresspeople for some similar forms of transparency
(To be clear, I’m aware this takes you pretty far down the train to crazytown or a dystopic surveillance state, and there’s probably only like 10 people in the world currently who think this would be a good idea. Consider these points as thought experiments about what might be possible, as opposed to a recommendation to actually implement any of these.)
Objections to transparency
1. It empowers your enemies
True. But it empowers your friends as well. Truth is an asymmetric weapon; if what you are working on is good, you will have more friends than enemies. (If what you’re working on is not good, you should hope to learn about this fact.)
One microscopic example of this: openbook.fyi (a 2023 side project of my now-wife Rachel Weinberg) led EA critic Emile Torres notice and tweet about funding patterns, helping Eliezer update against Slime Mold Time Mold.
2. Sometimes it leads to worse outcomes
True. But I think a general policy of being transparent always will lead to better global outcomes, for AI safety (and, for the world).
Also, I think the argument for transparency comes from more virtue-ethics than consequentialist grounds. And there’s a consequentialist-y reason to engage with virtue ethics: virtue ethics works!
3. Infohazards exist
Okay, I might be willing to hear this out for things like “how to engineer novel pathogens”. But I think this is sometimes used as a justification for being, idk, lazy about what you publish and write.
4. Reputational hazards exist
(I’ll leave this one as an exercise for the reader.)
5. Transparency is costly
True. And I understand that this is a big part of why OpenPhil/CG dialed down its transparency. I just think that the benefit outweighs the costs.
Also transparency doesn’t have to be quite as costly, if you bake it into your operating principles from the start. It feels self-aggrandizing to harp on this, but: Manifund has been open source from day 1, with its proposals and grants always open to the public. We write about our actions, engage with public criticism, and aim to reason in public and update where we’re wrong.
6. AI safety is already pretty transparent
True! And, this is again part of what I loved about the movement. I wish to see this transparency continued and amplified.
Even if AI safety is already doing better on transparency than its reference class (other nonprofits, other researchers, other political actors, labs), I think AI safety must take it upon itself to be a paragon of virtue, given its ambitions.
I am very big on transparency, I strongly agree with this stance! (Admittedly it can be hard on a practical level sometimes). I am thinking more about how charities can open their data up. For example, at Kaya Guides we have our live metrics posted to our website, and I have had a lot of people I wouldn’t normally hear from reach out to tell us how much they liked it. So I think there are also some positive gains to be made there!
I would also note that another objection to transparency is regulatory hazards. This is quite common in the global health field. Sometimes it can be hard to justify transparency today in case a future risk becomes apparent and you might retroactively regret transparency.
Thanks! Cool to see Kaya's data; I've always been proud of pushing Manifold to make its stats page public as well
Can you say more about regulatory hazards in global health, and/or give an example? I'm imagining something like "you publish what you did in a developing country, then the regime changes, and then they notice you did a bunch of things they don't like, and put you in jail"