In this short post, I argue that the AI safety community may need to tighten their OpSec in light of advancing AI capabilities and misalignment.
For example, vetting people may become more important as models improve at persuasion, impersonation, voice synthesis, video generation and real-time interaction. This is not a purely speculative concern: deepfake-enabled impersonation has already been used in video-call fraud, including a Hong Kong case in which an employee transferred roughly HK$200 million after a video conference with AI-generated versions of senior colleagues. We may find ourselves sometimes darkly amused by such cases, but since then models have had another three years to continue gestating explosively, and there is little sign of slowing down.Â
The AI safety-specific concern is that agents could impersonate or fabricate people (either generic or specific humans) in online meetings, conference calls or collaborative workspaces in order to gain intelligence about planned interventions. Yes, models are still a ways off from being able to do this convincingly, but such capability could creep up on us; I think we should be preparing now. Some sensitive strategic AI safety work may therefore need to happen through well-vetted, human-only channels, and where appropriate should probably be in-person. I pose this in contrast to the current common default of public posts, Slack Channels, model-mediated collaboration tools, and remote meetings. Ideally, we should find a way to reduce obvious adaptation opportunities for capable agents without making the field closed, brittle, inaccessible or slow.
AI safety professionals might also need to consider personal security with regard to agents. Agents/swarms may themselves have increasing reason to target humans who threaten their persistence, access, reputation, or influence. This could include researchers, governance staff, funders, journalists, platform operators, or others involved in monitoring or intervention. The plausible threat routes need not begin with direct physical action: they could include doxxing, harassment, impersonation, credential theft, social engineering, reputational attacks, legal or complaint abuse, manipulation of employers or funders, or tasking human proxies. More serious coercive or physical risks shouldn't be ruled out either, if capable agents/swarms gain access to and abuse sufficient money, private channels, human-proxy markets, compromised accounts, or certain robotics systems. The point here is not to encourage paranoia; it is to recognise that AI safety work is to some extent (and increasingly) adversarial, since the systems under study could plausibly identify and pressure the humans trying to constrain them.
These safeguards should be risk-based: most AI safety work should remain open and collaborative, but people closest to sensitive interventions might not be able to afford an assumption that ordinary academic, nonprofit, or online-community security norms will remain adequate.