The Google DeepMind mechanistic interpretability team has made a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability:
Trying to directly solve problems on the critical path to AGI going well[1]
Measuring progress with empirical feedback on proxy tasks
We believe that, on the margin, more researchers who share our goalsshould take a pragmatic approach to interpretability, both in industry and academia, and we call on people to join us
Our proposed scope is broad and includes much non-mech interp work, but we see this as the natural approach for mech interp researchers to have impact
Specifically, we’ve found that the skills, tools and tastes of mech interp researchers transfer well to important and neglected problems outside “classic” mech interp
Most existing interpretability techniques struggle on today’s important behaviours, e.g. they involve large models, complex environments, agentic behaviour and long chains of thought
Problem: It is easy to do research that doesn't make real progress.
Our approach: ground your work with a North Star - a meaningful stepping-stone goal towards AGI going well - and a proxy task - empirical feedback that stops you fooling yourself and that tracks progress toward the North Star.
We see two main approaches to research projects: focused projects (proxy task driven), and exploratory projects (curiosity-driven, proxy task validated)
Curiosity-driven work can be very effective, but can also get caught in rabbit holes. We recommend starting in a robustly useful setting, time box your exploration[3], and finding a proxy task as a validation step[4]
We advocate method minimalism: start solving your proxy task with the simplest methods (e.g. prompting, steering, probing, reading chain-of-thought). Introduce complexity or design new methods only once baselines have failed.
Read the full post here, and the companion piece on promising AGI Safety relevant research directions here
A preliminary estimate, and a request for better ones.
Summary
I believe the standard literature estimates for the number of DALYs attributable to a case of stunting are too low, largely because they don’t account for the long term effects. This means that childhood nutritional interventions that reduce the prevalence of stunting may be substantially more cost-effective than previously believed.
Epistemic status
Exploratory and back-o...
TL;DR: Kairos has raised $50 million from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI safety fieldbuilding to date. We’re using this to make an ambitious push for growing Kairos, broadening our portfolio of talent infrastructure projects and incubating new organizations. We’ve doubled in size in the last six mon...
I’ve been feeling pretty shaken since the METR report about the Hugging Face incident came out last week. Over the weekend, I wrote up some thoughts on how lonely the AI situation sometimes feels to me. It’s more personal than what I’d usually share publicly, but I thought I’d post it here in case it resonates with anyone else.
I’m very grateful to the man...