Head of Impact at ProVeg International, where I lead monitoring, evaluation and learning across six impact areas and 14 country offices. Behavioural scientist by training (PhD) and previously Deputy Director of Experimental Research at Yale, directing experimental behavioural research across the US, UK and India, and Senior Advisor to DEFRA's Strategic Behavioural Insights Team.
My research sits at the intersection of behaviour change and systems change: demand reduction in the illegal wildlife trade, climate communication and activism, and the social diffusion of plant-based diets. ~1,400 citations, h-index 20. Mostly I care about whether interventions actually work, and whether anyone bothered to check.
Nice to see the causal chain out laid out!
The counterfactual impact of (1) is that individuals "do good" they otherwise wouldn't have, but not necessarily that more good is done in the world overall. That relies on the pool being talent-constrained, so jobs would either sit empty, or be filled by demonstrably less effective candidates.
I would lean to the employer-side evaluation. Might be the costly option, but it's more informative, and could potentially be done via asking hiring managers at the point of hire, who else was in the final set and what they would have done if this candidate had not applied. That's close to a one-question survey.
Also worth surveying hiring managers on what are the biggest blockers - e.g., contextual knowledge? Specific skill sets? Could be useful for lighter touch evaluations of the programme down the road (i.e., are participants actually leaving with better contextual knowledge?)
I'm less sure about putting a money value on months saved early on - quite a few assumptions would be involved.
I would definitely recommend pre-registering the outcome definitions and the analysis before anyone sees the data, timestamped somewhere public. Ideally the outcome coding would also be done by someone blind whether participants were part of the cohort or not.
I think the wartime framing actually argues for the retrospective design - you wouldn't need to randomise anything, ask a mentor to take a fellow they did not choose, or make a single cohort worse.
Ah, very fair point! Screening is a real service to employers, and even if programmes produced no counterfactual placements at all they may still save the ecosystem a lot of search cost. It is measurable too, at least crudely, by asking hiring managers what they would have spent finding the same person.
I keep coming back to this question too. The whole case for spending heavily on talent pipelines rests on the field being talent constrained rather than absorption constrained, and I haven't seen public tests of that assumption. I think it would make another super interesting study, and is very testable! Vacancy durations, re-advertisement rates, applicant to hire ratios by role type, and hiring managers asked directly why a search failed would get you a long way.
I'm tempted to look at this properly as a follow up, so if anyone reading has hiring data they would be willing to share, or strong priors in either direction, I'd love to hear from you!
Thanks Ryan, that's useful context on the acceleration framing, and it actually might make a threshold design easier! Time to first role is continuous and observable, so your 10% and 14% figures could be compared to the "untreated" group.
It's also great to know that Coefficient's has an unpublished internal analysis. I do think making it public would be helpful so programmes can learn from each other. (And so MEL nerds like me can read it!)
Big fan of equivalence testing, agreed it would be important here!
I lead the Impact team at ProVeg International, and it's definitely unlocked our capabilities! E.g., one of our country office's recently developed a 3 year plan, and asked if we could forecast their impact in terms of GHG emissions. I reckon using Claude saved us close to a week in building the model.
Thanks, great questions!
RE long tails: I'd push back slightly here, because I think the tail argument actually cuts the other way. If impact is heavily long-tailed and the tail sits with the most impressive applicants, those are also the people most likely to have found their way into an AI safety role without the fellowship. The slam-dunk applicant has high absolute impact, but low counterfactual impact. So the margin may be where the programme's causal contribution is largest, even though expected impact per person there is smaller.
But I agree the RDD doesn't give us an average treatment effect across the cohort. It gives us a local estimate at the threshold. This is probably the most decision-relevant number for funders, (e.g., helps answer whether to accept 70 people instead of 60), but not necessarily the best number for "should this programme exist at all".
RE measuring outcomes: I think each programme's theory of change has to come first here. I've assumed it is roughly getting talented people into AI safety roles in a timely manner. (MATS reports the share of alumni "working in AI alignment", LASR reports 90% "gone on to work in AI safety/security", BlueDot reports graduates landing "impactful roles".) Public observation (e.g., LinkedIn) gets us a long way for this.
If a programme's real theory of change is latent capability that would never show up in a career trajectory, this does become more difficult!
RE power: completely agree. This is the main reason I think we need a joined-up approach. Pooling across programmes should give us a decent sample size (power calculations still depending), though at the cost being able to isolate any individual programme.
Re: Value Correlation - this is something we could feasibly actually check rather than just argue about. Conservation and development both have decades of interventions with measured short-run outcomes and later follow-up. I wonder if anyone has tried to estimate the correlation empirically?