This is a much-needed reality check. Sitting on ready-made near-miss data in Airtables and never running a regression discontinuity design shows how much low-hanging fruit is being ignored.
Most of these programs are just measuring their selection filter. When 30% say they’d do it unpaid, you aren't building a pipeline—you’re just stamping people who would have broken into the field anyway.
Hopefully funders like Coefficient Giving push for real evaluation standards. Dropping $50M on shallow retrospectives while skipping basic counterfactuals is wild for a movement built on rigor.
I really like that you are digging into this! 2 quick questions:
I could imagine the impact of talent interventions could be quite long tailed (e.g. perhaps 1 person in a cohort of 60 contributes ~half of the impact), might that make your proposed RDD approach understate impact, since it really just tells you the impact on the marginal accepted applicant (rather than the super impressive applicant who was a slam dunk)?
I agree that the application scores (assume that is indeed a thing) make the RDD approach straightforward, but how would you measure impact post-program, especially for the comparison group that might be hard to follow up with? I'm imagining here that, say, an AIS bootcamp greatly increases someone's AIS knowledge but not in a legible way that would show up in, say, their resume -- developing an assessment tool that captures this impact might be difficult (but also not something I've given serious thought to, and perhaps you have!)
Lastly just a statistical power flag: you link to LawrenceC's piece on variation in application assessments. That variation would suggest we'd need quite a good sample size to allow for precision in the detectable effect, which might be hard. I guess that's just a note to make sure some sample calculation is done once you have clarity on your outcome of interest to make sure you have a path to measuring the kind of effect that would make these programs cost-effective. The worst case scenario would be a study with 100 people looking to detect an impact of, say, 0.05 SD, when you are in fact only powered to detect 0.3 SD and basically doomed to concluding the program doesn't work, when perhaps it does.
These are intended as genuine questions to see if your ideas could work, if there is indeed a path to a good study here I'd love to see something proposed to, say, the EA Infrastructure Fund, your ideas seem valuable! (Disclaimer, I work at CEA, which contains the EAIF, but I am not involved in EAIF decisions and these are personal takes)
RE long tails: I'd push back slightly here, because I think the tail argument actually cuts the other way. If impact is heavily long-tailed and the tail sits with the most impressive applicants, those are also the people most likely to have found their way into an AI safety role without the fellowship. The slam-dunk applicant has high absolute impact, but low counterfactual impact. So the margin may be where the programme's causal contribution is largest, even though expected impact per person there is smaller.
But I agree the RDD doesn't give us an average treatment effect across the cohort. It gives us a local estimate at the threshold. This is probably the most decision-relevant number for funders, (e.g., helps answer whether to accept 70 people instead of 60), but not necessarily the best number for "should this programme exist at all".
RE measuring outcomes: I think each programme's theory of change has to come first here. I've assumed it is roughly getting talented people into AI safety roles in a timely manner. (MATS reports the share of alumni "working in AI alignment", LASR reports 90% "gone on to work in AI safety/security", BlueDot reports graduates landing "impactful roles".) Public observation (e.g., LinkedIn) gets us a long way for this.
If a programme's real theory of change is latent capability that would never show up in a career trajectory, this does become more difficult!
RE power: completely agree. This is the main reason I think we need a joined-up approach. Pooling across programmes should give us a decent sample size (power calculations still depending), though at the cost being able to isolate any individual programme.
when you are in fact only powered to detect 0.3 SD and basically doomed to concluding the program doesn't work, when perhaps it does.
That’s not what you should conclude from a failure to reject the null at (say) an 0.05 significance level. You would need an equivalence test to ~conclude a program (probably) does not work.
You can make meaningful statistical inference for decision making with a lower level of statistical power, especially within a Bayesian framework.
Nice! Though I guess my concern that someone would conclude the program hadn't work, after failing to reject the null, is still there. But if the proposed study goes ahead, I'm sure your quick input could be very helpful!
Headline finding: I audited 17 AI Safety Talent programmes. Zero of 17 have published any comparison group, rejected-applicant follow-up, matched control or randomisation. Not one. Every programme that mentions a counterfactual does it by asking participants to self-report.
Background
At least $70 million[1] has gone into AI safety talent programmes so far, but we still can’t really say much about their impact. Notably, Kairos has just raised $50M on a self-described "shallow retrospective" that looked at "around 80 people". I decided to do a deep dive into how AI Safety Talent programmes self-evaluate their impact.
Disclaimer, these programmes are small, fast, and generally run by people with no evaluation training and no slack. The oldest one was only established in 2018 (AI Safety Camp, as far as I can tell). Most last less than 6 months. When timelines are genuinely short and urgent, then time spent evaluating is time that could be spent building. I am not criticising any individual programme.
Having said that, there are retrospective evaluation designs that could be run cheaply without slowing anything down. And a field that funds talent pipelines at this scale without knowing whether they work is not moving fast, it is moving blind.
A quick word on who I am: I am an evaluation expert, but I am new to AI safety. I have no stake in any of the programmes below and I am not an alumnus of any of them.
The audit
I audited 17 AI Safety Talent programmes: MATS, ARENA, LASR Labs, PIBBSS, SPAR, Pathfinder, Global Challenges Project, Kairos (org level), BlueDot Impact, AI Safety Camp, Apart Research, Talos Fellowship, Astra (Constellation), Anthropic Fellows, MARS (CAISH), ERA Cambridge, Horizon Fellowship. I used AI to initially search and extract findings, then checked every figure against the live page. Full findings here[2].
Out of 17 programmes:
Publish any comparison group, control or rejected-applicant follow-up
0
Publish a response rate for the survey behind their outcome claims
2
Publish a fully loaded cost per outcome
1
Pre-specified their outcome metrics before a cohort ran
0
Publish methodological caveats alongside their headline figure
2
(Because it bears repeating) Zero of 17 have published any comparison group, rejected-applicant follow-up, matched control or randomisation. Not one. Every programme that mentions a counterfactual does it by asking participants to self-report. And every one of them rejects a large majority of applicants, scores them numerically, and holds their contact details, so each is sitting on a ready-made near-miss comparison group it has never used.
Only two of the 17 publish a response rate for the survey behind their outcome claims. Everyone else with a survey publishes a percentage without telling you how many people it came from.
Rigour appears to run inversely to funding. AI Safety Camp operates on $60,000 to $300,000 a year, has never held a major-funder grant, and is the only programme here to have commissioned an external assessment and published it in full (including the estimate least flattering to itself). BlueDot Impact has raised $35 million and publishes "25% of our graduates land impactful roles within six months" with no definition of impactful roles, no denominator and no method. This appeared to be a common pattern: the organisations with the most money shared the fewest details.
The strongest numbers are presented in the weakest documents. LASR Labs's "90% of alumni have gone on to work in AI safety/security" appears only in recruitment posts, while Talos's "70% of our alumni (58/83)" appears only in a donation appeal. Neither organisation has published a retrospective or an impact report at all. The programmes that do publish reports (like MATS and ARENA), report more modest figures alongside limitations.
Multiple organisations contradict themselves across their comms. I won’t name names here, but figures on how many fellows get a job offer can differ between blogs and job postings, or placement rates can differ between accomplishments and recruitment pages. I’m sure there's nothing malicious here, I know from experience with a big org it’s easy for multiple figures to exist as evidence is updated! It does make it harder to do fair comparisons however.
Thoughts on the counterfactual impact of talent programmes
The counterfactual impact is what would have happened if your intervention didn’t exist. Chris Clay assembled the largest alumni dataset in the field, 600+ profiles across nine fellowships, and labelled counterfactual impact "The Golden Egg". He asked publicly for people to work on "What proportion of people would have entered AI Safety without doing a given fellowship?" That was January 2024, and as far as I can tell nobody has.
MATS was one of the most transparent programmes in this audit. They publish a 46% response rate alumni survey[3] with "possibly" and "probably" pooled into one counterfactual category, and no comparison group. In their report is possibly the most significant finding to me: 30% of alumni say they would have participated for no stipend at all. People motivated enough to do a research fellowship unpaid are precisely the people most likely to have found their way into alignment research anyway. The programme may be measuring its selection function and reporting it as its treatment effect.
This echoes a question that frequently occurs to me about career programmes. Unless a field is significantly talent constrained/requires very specific hard-to-get skills, it’s unclear to me how valuable the counterfactual impact of that programme was. For example, say Org A is hiring a researcher, while Org B has a career programme. Org B runs a fellowship/accelerator/has 1-2-1 coaching sessions, and an alumnus gets the job at Org A. The counterfactual impact relies on both 1) would the alumnus have applied for and gotten the job without Org B?, and 2) would Org A have been able to hire another person who is equally successful in the role?
A placement is only field growth if the person would not otherwise have got there and the role would not otherwise have been filled by someone equally good. If either leg fails, the programme has moved a person around rather than adding one. Self-report surveys, regression discontinuity designs etc, attempt to answer the first question, but none of the programmes audited have published anything on the second.
In fairness, the second question is harder, because answering it means asking employers rather than participants. 80,000 Hours is the only organisation I found that goes near it, noting of its headhunting placements that "these 15 cases involved an organisation hiring a candidate who they weren't otherwise exploring".
Whether that matters depends on where the binding constraint sits. Weronika Żurek's review of talent constraints suggests the AI Safety field is short of senior mentors, experienced operators, fieldbuilders and grantmakers, and she estimates 2,000 to 2,500 fellows will graduate from research programmes in 2026 alone. If that is roughly right, junior research talent is the part of the pipeline least likely to be constrained, and a placement rate from relevant talent programmes is less relevant.
To be clear, getting the right person into the right role faster is worth money, and so is the counterfactual for someone who would never have found the field at all. But those are different claims with different evidence requirements, and a placement percentage cannot tell them apart.
Jamie Bernardi, who ran AI Safety Fundamentals at BlueDot, wrote in 2024 that "people take your course because they're already interested in the topic – so you'll never know the counterfactual of which steps they took because of your course." The external assessment of AI Safety Camp says the same thing about that programme's own headline number: "AISC draws from people already interested in AIS, so researchers who appear to have a step-change in participation may not have needed AISC to break into AIS research in the first place."
Proposed design
Regression discontinuity design (RDD) is a quasi-experimental pretest–posttest design. It uses a specific cutoff score on a continuous assignment variable to estimate the causal effect of an intervention. In the case of AI Safety Talent Programmes, we know that they are massively oversubscribed and the cutoff for being accepted is relatively arbitrary. Applicants who were close to being accepted but didn’t quite make it are probably not materially poorer researchers than those who did just make the grade.
LawrenceC collected a senior researcher's retrospective assessment of 60 junior AI safety researchers. The mean change between initial and later ratings of promise was 0.05 on a ten-point scale with a standard deviation of 1.37, and the two strongest performers had initially been rated "somewhat above average" and "below average". Żurek puts it more plainly: "even the best candidate assessment methods are just moderately correlated with job performance." That doesn’t mean selection is worthless, it just means the ordering is informative in the large and close to a coin toss within a few points of the threshold (exactly the condition a discontinuity design needs).
A RDD design first needs to define what success looks like. Working full-time on AI safety at eighteen months post-decision is an obvious primary, with publication as a secondary. Both are observable from the outside and do not depend on anyone answering a survey.
The minimum viable version is retrospective. Application scores already exist in an Airtable at every one of these organisations. Pull the scores, pull (or search[4])the outcomes for admits and near-misses, done. No surveys needed, no cohort disrupted.
(Quick caveat, application scores were collected to make a selection decision, and studying career outcomes is a different purpose, so a programme shouldn’t simply repurpose them. It needs either fresh consent from past applicants or a documented legitimate-interests assessment, with pseudonymised analysis and no individual reported. Going forward, I’d recommend a checkbox at the application stage for the use of data for research purposes.)
Now, it’s worth noting that these cohorts are small. MATS has 60 scholars, LASR 12 to 16, ARENA 30. Regression discontinuity needs a decent density of applicants near the threshold, so most of these programmes would be individually underpowered. Ideally we would need a pooled analysis across programmes. Failing that, we’d need to simply compare admits against near-misses without the discontinuity machinery.
In the academic literature regression discontinuity is a standard design. E.g., see Bol, de Vaan and van de Rijt (2018, PNAS) on early-career grant thresholds (a near-miss design on a competitive fellowship, which is structurally what these programmes are), or Gonzalez-Uribe and Leatherbee (2018, RFS) on a start-up accelerator (a competitive cohort programme with numeric application scores).
In the EA field, the closest example I can find is Fodor & Tidmarsh (2024), who conducted a case-control survey of EAGx attendees with a difference-in-differences design. Note they had an admissions process available as an assignment mechanism but did not use it.
Conclusion
In a recent retrospective on their Effective Careers RFP, Coefficient Giving said "we saw a wide range of how applicants monitor their outcomes, and we're considering an internal project to help set and share clearer standards." Given the small cohort numbers, I agree that a joined up approach across the field, with standardised metrics, would be valuable. Failing that, it would be great to see one of the bigger organisations take this seriously. I’m happy to help design a study, or put people in touch with interested researchers.
I know filling out surveys is annoying but c’mon guys, if you get given $$$ to study something cool, do the decent thing and fill out a form for your organisation after.
Less onerous than it sounds, because the outcome side is largely public. Whether someone works at a safety organisation, and whether they have published, is visible on LinkedIn and in conference proceedings. Kairos is already building Talent Commons as a "consent-first database for sharing information about program participants and job applicants", which is most of the infrastructure this would need.
TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/
Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don't realize contradict each other, or that have strange implications that they wouldn't endorse if they thought about it more.
Not just that—population ethics is a BIG DEAL. A lot of ethical decisions hinge on how you think about changes in future populations....
A preliminary estimate, and a request for better ones.
Summary
I believe the standard literature estimates for the number of DALYs attributable to a case of stunting are too low, largely because they don’t account for the long term effects. This means that childhood nutritional interventions that reduce the prevalence of stunting may be substantially more cost-effective than previously believed.
Epistemic status
Exploratory and back-o...
TL;DR: Kairos has raised $50 million from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI safety fieldbuilding to date. We’re using this to make an ambitious push for growing Kairos, broadening our portfolio of talent infrastructure projects and incubating new organizations. We’ve doubled in size in the last six mon...
Why don’t funders require randomisation? Eg pick the 2n best applications, randomly admit the n best, track outcomes etc.