There are education programs in low-income countries that deliver the equivalent of three years of high-quality schooling for about $100 per child. GiveWell doesn't fund them.
We can see the effects of education on an individual level too:
So why doesn’t GiveWell fund educational interventions? We have little idea how much those learning improvements will lead to better lives on average. It’s time-consuming and expensive to conduct an education experiment that follows kids from a young age all the way into their adult lives (to measure their income and health). As a result, there are only a handful of developing-world RCTs linking education to adult incomes. None of them can isolate learning as the cause, and only one (a scholarship RCT in Ghana) is designed to even get close.
It makes sense to be conservative in the face of limited evidence, but does it make sense to place no value on test score gains or increased school years, as GiveWell does?
My main takeaways:
Our current understanding is that there is limited high quality evidence establishing a causal link between increased time in school or test scores to improvements in life outcomes such as earnings or health, which we consider to be better measurements of general wellbeing.
There is no gold standard RCT evidence that links developing world learning improvements to income. However, there is a lot of less direct and lower quality evidence to weigh. Albinsky’s 2023 post digs into this work. Most of these studies don't report a test score -> income effect directly. He derives an effect from each paper, and the straightforwardness of the derivations varies. He estimates:
What do we do with all that imperfect evidence?
One option is to not even attempt to make an estimate, like GiveWell does. Another is to more or less accept it as is and not worry too much about causality, which economist Lant Pritchett proposes:
I don't feel one can be "neutral" or just "bracket" this by saying "Hey, I am just going to ignore those studies because there is a story in which the association "might" be due to this or that and not causal.
I’d argue the best approach is similar to Albinsky’s: make the most of what we have! These studies are not perfect, but there are a lot of them, and we also have some intuition/priors. An organization like GiveWell is in an even better place to build informed intuitions given their experience in the developing world.
That said, this evidence has clear weaknesses:
Albinsky argues that each piece has flaws, but together they amount to a fairly robust literature, with consistent effects across all of them. He combines the various estimates above by averaging the effect in each of the 5 buckets into a 19% income increase per test score SD increase estimate.
He then applies a validity discount modeled on the adjustments GiveWell makes to its own deworming and malaria income estimates. A whole post could be written on all the assumptions going into these discounts, but let’s trust Albinsky on this point and focus elsewhere. A 17.5% multiplier (FYI GiveWell applies a 13% multiplier to deworming) brings him to a final estimate of a 3.3% income rise per SD increase in test scores.
Albinsky made his case, which got a lot of attention on the EA Forum, but no public response from GiveWell. Has there been any research in the three years since his piece that we should update on?
Ghana/Duflo
There is an important recent update to the best LMIC scholarship study, an RCT in Ghana. Duflo et al. 2024 looks at the children of scholarship recipients and finds a 45% reduction in under-3 mortality (p = 0.065) and a 0.24 SD increase in cognitive development at age 5 (p = 0.005) and 0.25 at age 7 (p = 0.035)! This effect is only for female recipients, but it's still a very promising finding. Those cognitive effects are even larger than the effects on the actual scholarship recipients.
Here’s a fun video on the paper:
The income findings from Duflo et al. 2026 (as of 9/6; this is a live working paper) are also promising. In the 2021 results referenced in Albinsky’s post, incomes were up 2.5% 11 years after the program started and participants started high school. Now it’s been 15 years, and incomes are up 11%, a four-fold jump in effect size. The p-value is 0.08 and the 95% confidence interval is -1% to 23%, so we are still looking at a wide range.
The bad news is that roughly half of the income gains come from winning public sector jobs, which are generally zero sum. Firms expand with more productive workers, whereas civil service just gets a longer queue. Someone getting a public sector job means someone else is not getting that job. Public sector jobs are typically paid above the market wage, which means this ends up being more of a transfer to scholarship recipients than a true increase in overall output and wealth.
There is little reason to expect the public sector effect to be unique to this study. This should make us concerned about the other voucher and scholarship studies, which, in my opinion, are the strongest lines of evidence. The authors of the Colombia paper expressed this concern, but were unable to measure it.
So how should we adjust Albinsky's estimate? Two things moved, in opposite directions and for different reasons. The income effect quadrupled over a longer follow-up time. The public-sector finding cuts the effect in half. Those point toward a potentially higher estimate than Albinsky’s, but I'll keep using 3.3% for the rest of this piece, because I don't feel confident enough to change it. That's my bigger takeaway. The best-designed study in this literature is unstable in both directions at once, and it's still a live working paper. Whatever number we use today could be stale by the next revision.
Publication Bias
Clark & Nielsen (2026) conducts a meta-analysis on “The Returns to Education”. They look at 53 compulsory education papers and find an average rise in incomes of 8.5% per year of schooling. However, after adjusting for publication bias, they find an effect of somewhere between 0% and 3%. Publication bias is partially dealt with through Albinsky’s GiveWell-esque validity discounting, but the findings from Clark suggest bias may be even higher than expected.
Patrinos & Psacharopoulos (2026) originally contained a rebuttal to the Clark paper. They argue we should not expect to see a normal distribution of effect sizes in a bias-free literature, because additional schooling is unlikely to actively cause harm and lead to a reduction in educational outcomes. This rebuttal was removed from the published version of the paper, but it's not clear why. There is some merit to the distributional argument they make and addressing it would probably move Clark’s estimate up a couple percentage points.
Where does all that leave us? We have two results pointing towards larger impacts: Duflo’s most recent income effects and their intergenerational effects. And we have two results pointing towards smaller impacts: Duflo’s zero-sum findings and Clark’s publication bias. I’d advocate for Duflo’s work to be the main focus of a CEA, and I think the chances of publication bias there are quite small (they would be reporting income effects even if they were null). For the rest of this piece I'll use Albinsky’s 3.3% and see where that gets us.
A couple examples are useful to understand what a 3.3% increase actually means in terms of educational intervention effectiveness.
The genesis for this piece was my analysis on glasses for workers. As I started looking into the effects of glasses for students, I stumbled into this methodological debate. So is it cost-effective to give glasses to students under Albinsky’s assumptions?
There have been a number of good (but geographically specific) RCTs looking at the effects of giving myopic students glasses, all finding that glasses led to test score increases between 0.11 and 0.25 SDs. Let’s optimistically assume a program cost of $40 per recipient (for screening and free glasses) and a 0.25 SD test score increase. With the 3.3% income gain per test score SD increase (plus assuming 40 years of increased income and standard GiveWell assumptions), we only get returns that are 2.8x over GiveWell’s benchmark, which does not clear their 6x bar.
The glasses case isn’t great, but what about something with stronger cost-effectiveness evidence like Structured Pedagogy?
Structured Pedagogy is a broad term for things like scripted/structured lesson plans, linked student/teacher materials, teacher training, and ongoing coaching/monitoring. One meta-analysis on structured pedagogy programs in Sub-Saharan Africa found average test score improvements of 0.23 SDs (when restricting to the strongest studies). Another meta-analysis, with little overlap and more broadly looking at LMIC programs, found a 0.14 SD improvement.
In a conservative scenario, let’s take the 0.14 SD estimate and assume a program will cost $15 per student. In an optimistic scenario, let’s take the 0.23 SD estimate and assume an $8 cost per student. With the 3.3% income gain per test score SD increase (plus assuming 40 years of increased income and standard GiveWell assumptions), that gets us to returns that are 4-13x over GiveWell’s benchmark. Even the conservative estimate is close to the bar, and this is only looking at direct income effects.
What about some of the untouched and potential large effects:
It's unclear how well these effects would extend to different intervention types, but if we throw some of them into the mix, already promising educational interventions start to look very promising.
We’re left with a lot of uncertainty. The best study in this literature, in a working paper that's still being revised, had a 4x change in income gains from one follow-up to the next. They also introduced a complication (zero-sum public sector jobs) that could halve our effect sizes. On top of that, we have several other underexplored mechanisms that could outweigh our income effects. We know education matters, but we don’t know how much, and that's a challenging environment for CEAs.
There is some great news, though. In 2024, GiveWell and Coefficient Giving funded a study scoping “the feasibility of conducting a long-term follow-up to an RCT of a low-cost preschool program implemented by Save the Children in Mozambique from 2008 to 2010”, and determined this was feasible. The follow-up study has begun, and they are aiming to finish data collection by mid-2027. Their goal is to get a follow-up rate of at least 85%, which would hopefully be enough to assuage selection bias concerns.
As they note, this study may not generalize and the sample size is similar to the Ghana/Duflo study that had wide confidence intervals. The cohort will also only be 21-23 years old, well before their peak earning years. The survey document doesn’t mention looking into public sector employment and zero-sum concerns, but I hope that will make it into the study.
There is also the Return to Learning Initiative from the Center for Global Development, which has identified 20 other education LMIC RCTs that are good candidates for long-term followups. Perhaps this could be sped up with more funding? I can’t track down any funding information.
I also wonder if there is room for more qualitative evidence. I think I have a decent sense for how similar people with different levels of education might fare in the US, and that's a useful background/prior for interpreting quantitative evidence and deciding what specific interventions to consider. GiveWell and Coefficient Giving have the resources to cheaply gain knowledge on what that looks like in developing countries. A few questions I’m curious about:
I can understand not accepting Albinsky’s estimates, and I’m glad GiveWell is pursuing more evidence. But I’d love to see more public information on their strategy (and others would as well! Kirsty Newman had a post about education last week). Are they planning to conduct CEAs on various educational interventions as the Mozambique and CGD results roll in? How do they decide how much money to invest in evidence generation and are they planning to invest more in educational evidence generation?