AI safety
AI safety
Studying and reducing the existential risks posed by advanced artificial intelligence

Quick takes

3
15h
TLDR: I've updated towards pausing further AI development indefinitely. When Scott Alexander proposed regulating AI like clinical drugs, many (including myself) balked at this given the sclerosis of FDA or related bodies. Yet HF updates me towards treating frontier AIs as nuclear. There onerous regulation is likely good. In general, most bureaucratised regulatory regimes are bad. In some cases, like nuclear technology or ensuring planes are safe to fly on, they're welfare improving. It's clear that AI belongs in the latter category. Also trace inversion is a thing, so you can get open source weights to be roughly similar to that of Fable/Mythos. The slowdown camp were right. At bare minimum, all labs should pause training, further development, and releases indefinitely until we figure out how to align AIs and regulate them (and their use by humans). I reckon the current capabilities we have now are sufficient for accelerating progress towards curing cancers etc. So my balance has shifted towards minimising the existential risks now. Pause AI advocates are correct. If governments could coordinate internationally to achieve such (big if), I'd support it. I used to be highly sceptical of doomer arguments, yet Hugging Face is almost a textbook LW scenario and no one knows how to spot or prevent such scheming. Sandboxing, guardrails, constitutions etc. don't work.
9
1d
How impactful would it be to copies of @Garrison's new book Obsolete to elected officials who belong to its political target audience (Dems, particularly left ones) and might not have been responsive to traditional x-risk-centric comms (e.g. IABIED)?
3
4d
[central europe] We are organising the largest AI-safety march (at least in central europe) to date at Prague, starting at 15:00 on 31st August. We have got parlimentary support as well as very supportive police, allowing us to take the preffered route. We want as many people to come - fell free to spread information about the event outside EA as much as possible, especially if you have friends living near Prague who may come. If you want closer info/visual materials or help us with organisation, you can get in contact with czech PauseAI on whatsapp. Luma event link  
1
9d
Most AI-geopolitics analysis concentrates on frontier production. I've been exploring a complementary question: what happens to the capabilities the frontier leaves behind? Downloadable models can persist after commercial support and meaningful developer control disappear. As inference and adaptation costs decline, some of these systems may become accessible to actors that could not have exploited them at release. The longer piece develops this through the ideas of capability sedimentation, temporal capability arbitrage and “compute-poor but model-rich” states. I'm especially interested in whether this should affect how the AI-governance community thinks about model-release decisions and proliferation below the frontier. https://fabiofrettoli.substack.com/p/the-geopolitics-of-model-graveyards?r=sb1v&utm_campaign=post-expanded-share&utm_medium=web 
15
13d
Applications to SPAR Fall 2026 are closing tomorrow Aug 18 EOD Anywhere on Earth. SPAR is the ecosystem's biggest AI safety research program, and it's part-time remote. We still have many strong projects across AI safety, AI policy, and biosecurity with very few applications[1], so please consider applying! This round, we also have 18 non-research/generalist projects that people can apply to, and we have much more mentee capacity than previous rounds; we expect to accept around 500 people into the program. If you have friends who have thought about going into AI safety, spread the word! 1. ^ To be specific, as of 2:25 PM PT, we had around 67 projects with fewer than 20 applications total!
7
2mo
Apply to be a Teacher for TARA Round 2, 2026 🧑‍🏫 • Part-time role ($80 AUD/hr, ~13 hours/week). • Lead Saturday sessions and provide remote support. • Must have strong ML skills and completed most or all of the ARENA curriculum. • Applications close 8 August 2026, reviewing candidates on rolling basis. • Check out the role description and apply.
1
2mo
Just published a thesis on how Goal-Setting Theory applies to principal-motivated deceptive agents and secret loyalties. Looking for feedback. https://forum.effectivealtruism.org/posts/bxALZuqcf5BXgvpEt/ai-agents-with-a-specific-secret-loyalty-are-more-dangerous
21
2mo
2
Please list any new funding opportunities you can think of here on the Forum? I feel like we might already be in the early ramp-up to significantly more EA aligned funding. At the same time, the Forum's overview over funding opportunities feels like it is quickly getting outdated. I think as things move quickly, coordination might become looser and new promising interventions are identified, it is helpful for people to have a good overview over available funding sources and their priorities. I have heard on the grapevine there is already funding on several fronts that might not be very public. I am a little bit uncertain if perhaps it is better these sources remain anonymous. At the same time, I think there might be several promising EA projects that are not sufficiently visible to people influencing funding decisions.  Epistemic note: I am not listing these yet as I have not had time yet to verify how much they qualify as EA funding opportunities. Here are a few recent developments I am considering listing on the funding opportunities page, but would like someone that knows these funds better to list them: * Probably several I have missed - please list these here * OpenAI Foundation - AI Resilience * OpenAI Rosalind * The Launch Sequence (not a fund in itself, but plausibly one can treat this kind of like a funding source?) * Several of Renaissance Philanthropy's (RP) funds (several quite plugged in EAs do not even know of RP!) * Astralis Foundation
Load more (8/263)