TL;DR
NOVAH (No Violence At Home) was incubated by Charity Entrepreneurship (now Ambitious Impact) in 2024 to test a promising idea: preventing intimate partner violence through edutainment, in our case a serialised radio drama. Over the past two years we have produced and aired two seasons in Rwanda.
We are currently evaluating our second season through a randomized controlled trial with 2,400 couples in Rwanda in partnership wi...
TL;DR
* The Long-Term Future Fund is closing down, and EA Funds is launching the聽Transformative AI Fund with a new full-time team.
* The fund's primary focus is technical AI safety and AI governance (including post-AGI governance), as well as supporting fields such as field-building and forecasting. We'll also consider non-GCR implications of transformative AI such as flourishing futures and digital...
The current Long Term Future Fund (LTFF) fund managers and I have decided to step back from our work on the LTFF. Because we believe LTFF donors trusted the fund managers to ensure that the funds would be used in line with the purposes of their donation, we've decided the right move is to close the fund.
While LTFF is closing, note that EA Funds has launched a new fund...
Congrats to the prizewinners!
Folks thinking about corrigibility may also be interested in the paper "Human Control: Definitions and Algorithms", which I will be presenting at UAI next month. It argues that corrigibility is not quite what we need for a safety guarantee, and that (considering the simplified "shutdown" scenario), instead we should be shooting for "shutdown instructability".
Shutdown instructability has three parts. The first is 1) obedience - the AI follows an instruction to shut down. Rather than requiring the AI to abstain from manipulating the human, as corrigibility would traditionally require, we need the human to maintain 2) vigilance - to instruct shutdown when endangered. Finally, we need the AI to behave 3) cautiously, in that it is not taking risky actions (like juggling dynamite) that would cause a disaster to occur once it is shut down.
We think that vigilance (and shutdown instructability) is a better target than non-manipulation (and corrigibility) because:
Given all of this, it seems to us that in order for corrigibility to seem promising, we would need it to be argued in some greater detail that non-manipulation implies vigilance - that the AI refraining from intentionally manipulating the human would be adequate to ensure that the human can come to give adequate instructions.
Insofar as we can't come up with such justification, we should think more directly about how to achieve obedience (which needs a definition of "shutting down subagents"), vigilance (which requires the human to be able to know whether it will be harmed), and caution (which requires safe-exploration, in light of the human's unknown values).
Hope the above summary is interesting for people!