Head of Impact at ProVeg International, where I lead monitoring, evaluation and learning across six impact areas and 14 country offices. Behavioural scientist by training (PhD) and previously Deputy Director of Experimental Research at Yale, directing experimental behavioural research across the US, UK and India, and Senior Advisor to DEFRA's Strategic Behavioural Insights Team.
My research sits at the intersection of behaviour change and systems change: demand reduction in the illegal wildlife trade, climate communication and activism, and the social diffusion of plant-based diets. ~1,400 citations, h-index 20. Mostly I care about whether interventions actually work, and whether anyone bothered to check.
Ah, very fair point! Screening is a real service to employers, and even if programmes produced no counterfactual placements at all they may still save the ecosystem a lot of search cost. It is measurable too, at least crudely, by asking hiring managers what they would have spent finding the same person.
I keep coming back to this question too. The whole case for spending heavily on talent pipelines rests on the field being talent constrained rather than absorption constrained, and I haven't public tests of that assumption. I think it would make another super interesting study, and is very testable! Vacancy durations, re-advertisement rates, applicant to hire ratios by role type, and hiring managers asked directly why a search failed would get you a long way.
I'm tempted to look at this properly as a follow up, so if anyone reading has hiring data they would be willing to share, or strong priors in either direction, I'd love to hear from you!
Thanks Ryan, that's useful context on the acceleration framing, and it actually might make a threshold design easier! Time to first role is continuous and observable, so your 10% and 14% figures could be compared to the "untreated" group.
It's also great to know that Coefficient's has an unpublished internal analysis. I do think making it public would be helpful so programmes can learn from each other. (And so MEL nerds like me can read it!)
Big fan of equivalence testing, agreed it would be important here!
I lead the Impact team at ProVeg International, and it's definitely unlocked our capabilities! E.g., one of our country office's recently developed a 3 year plan, and asked if we could forecast their impact in terms of GHG emissions. I reckon using Claude saved us close to a week in building the model.
Thanks, great questions!
RE long tails: I'd push back slightly here, because I think the tail argument actually cuts the other way. If impact is heavily long-tailed and the tail sits with the most impressive applicants, those are also the people most likely to have found their way into an AI safety role without the fellowship. The slam-dunk applicant has high absolute impact, but low counterfactual impact. So the margin may be where the programme's causal contribution is largest, even though expected impact per person there is smaller.
But I agree the RDD doesn't give us an average treatment effect across the cohort. It gives us a local estimate at the threshold. This is probably the most decision-relevant number for funders, (e.g., helps answer whether to accept 70 people instead of 60), but not necessarily the best number for "should this programme exist at all".
RE measuring outcomes: I think each programme's theory of change has to come first here. I've assumed it is roughly getting talented people into AI safety roles in a timely manner. (MATS reports the share of alumni "working in AI alignment", LASR reports 90% "gone on to work in AI safety/security", BlueDot reports graduates landing "impactful roles".) Public observation (e.g., LinkedIn) gets us a long way for this.
If a programme's real theory of change is latent capability that would never show up in a career trajectory, this does become more difficult!
RE power: completely agree. This is the main reason I think we need a joined-up approach. Pooling across programmes should give us a decent sample size (power calculations still depending), though at the cost being able to isolate any individual programme.
Re: Value Correlation - this is something we could feasibly actually check rather than just argue about. Conservation and development both have decades of interventions with measured short-run outcomes and later follow-up. I wonder if anyone has tried to estimate the correlation empirically?
Love this!
One suggestion: you mentioned IPV prevalence in Rwanda is ~25%*, but you're only powered to detect a 2pp effect? 2pp reduction with a baseline 25% prevalence is an 8% relative reduction. Anything below that is unlikely to be significant. You're unable to reliably detect a 6 or 7% relative reduction which is above your 5% bar, but would still send you into "pause and explore". I might be better to commit on the confidence interval rather than on significance. Lakens' work on the smallest effect size of interest (SESOI) and equivalence testing could be useful here, since it would let a null result mean you have ruled out effects above some size rather than just failed to find one.
*across a year, while you're measuring across four months. Baseline prevalence is probably smaller in that time period, meaning the minimum detectable effect gets larger in relative terms.
Fully agree with most of what's written here!
>A strong monitoring plan includes a handful of indicators, with the project assigning each a probability of success. These can be scored once a year, with any failures left in. It should take the project around half a day to complete and track things it wants to know anyway.
Just a note on MEL from the inside of an org, it usually takes more than half a day, and that's just for one funder. It gets extremely time-consuming when you have multiple funders each negotiating their own indicator set. What about major donors coordinating on a shared indicator vocabulary?
"75% of organizations say improving MEL is important or urgent, but only 18% would definitely pay for it."
That's a rather upsetting finding. To turn MEL into an ongoing function rather than just a one-off deep dive you need structures in place, people with expertise, tools to collect data etc. It's worth the investment, but funders/orgs definitely have to be willing to put money behind it.
Looking forward to seeing what you produce!