Epistemic status: Speculation from two decently informed advocates armed with anecdata.
Note on process: After having some version of this conversation several times and saying, “we should probably write about this publicly,” we took the less heroic route: we recorded one of our conversations, fed the transcript into an LLM, and then substantially revised the structure, substance, and framing ourselves. We will not be sharing the transcript, as it is in...
TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/
Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don't realize contradict each other, or that have strange implications that they wouldn't endorse if they thought about it more.
Not just that—population ethics is a BIG DEAL. A lot of ethical decisions hinge on how you think about changes in future populations....
Headline finding: I audited 17 AI Safety Talent programmes. Zero of 17 have published any comparison group, rejected-applicant follow-up, matched control or randomisation. Not one. Every programme that mentions a counterfactual does it by asking participants to self-report.
Background
At least $70 million...
Anthropic's September report on misuse of its models poses interesting questions about AI safety. The report, detailing actions ranging from the use of Claude to set up a fake online dating profile farm to the model being employed by the Houthis to design software to guide missiles, was not produced under any legal obligation. Under the TFAIA in California and the EU AI Act, it is only mandated to confidentially disclose safety incidents to authorities, and the jury is still out on whether some of these cases would fall under the definition of 'serious incidents' (in the EU AI Act case) or 'critical safety incidents' (in the TFAIA case). Potential explanations for why the company still decided to produce such a report are not hard to imagine: better PR, internal employee pressure, and institutional inertia. However, these incentives are largely internal to the company and can change. If the board decides that these reports should be toned down to dissociate Anthropic from dangerous uses of AI (perhaps to prepare for an IPO), there is no reason why this would not take place. That would be a serious blow to the cause of AI safety at a time when the potential fallout from increasing model progress has still to be spelled out. Any loss of data on this front would mean one more threat vector left unexplored, and a part of society and the world left potentially unprepared.