TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/
Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don't realize contradict each other, or that have strange implications that they wouldn't endorse if they thought about it more.
Not just that—population ethics is a BIG DEAL. A lot of ethical decisions hinge on how you think about changes in future populations....
I’ve been feeling pretty shaken since the METR report about the Hugging Face incident came out last week. Over the weekend, I wrote up some thoughts on how lonely the AI situation sometimes feels to me. It’s more personal than what I’d usually share publicly, but I thought I’d post it here in case it resonates with anyone else.
I’m very grateful to the man...
Headline finding: I audited 17 AI Safety Talent programmes. Zero of 17 have published any comparison group, rejected-applicant follow-up, matched control or randomisation. Not one. Every programme that mentions a counterfactual does it by asking participants to self-report.
Background
At least $70 million...
A lot of AI safety people I know are excited about Anthropic and I don't fully understand why. My instinct is to distrust company because they're the one pushing the AI arms race, autonomously hacking three companies, and IPO'ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities?
This becomes much easier to explain when you allow for the possibility that people are biased or irrational or have conflicts of interest. People want to work on cool problems, they want to be close to the action, they want to get rich, and they want to think of their friends as good people; then they reason backward from that bottom line to determine that Anthropic must be the good guys.
Anthropic has produced a lot of alignment research. On certain theories* of where AI danger comes from, that research has been useful enough that Anthropic's overall impact is net positive.
*Those theories are wrong. In a sentence: Anthropic's alignment research is almost entirely centered around how to produce desired behaviors in the short term, with no understanding of how to make an ASI continue to be aligned once you are no longer smart enough to detect misalignment.
I don't think they're respected within this space? But of course 'this space' is vague.
By 'this space' I meant AI Safety. I at least see a lot of Anthropic roles being published on 80k and have seen AI safety clubs direct people towards Anthropic fellowships as they would MATS.