Anthropic's September report on misuse of its models poses interesting questions about AI safety. The report, detailing actions ranging from the use of Claude to set up a fake online dating profile farm to the model being employed by the Houthis to design software to guide missiles, was not produced under any legal obligation. Under the TFAIA in California and the EU AI Act, it is only mandated to confidentially disclose safety incidents to authorities, and the jury is still out on whether some of these cases would fall under the definition of 'serious incidents' (in the EU AI Act case) or 'critical safety incidents' (in the TFAIA case). Potential explanations for why the company still decided to produce such a report are not hard to imagine: better PR, internal employee pressure, and institutional inertia. However, these incentives are largely internal to the company and can change. If the board decides that these reports should be toned down to dissociate Anthropic from dangerous uses of AI (perhaps to prepare for an IPO), there is no reason why this would not take place. That would be a serious blow to the cause of AI safety at a time when the potential fallout from increasing model progress has still to be spelled out. Any loss of data on this front would mean one more threat vector left unexplored, and a part of society and the world left potentially unprepared.
I guess the discussion should not turn so much around shaming and blaming particular companies, but changing the structural incentives driving forward the current wave of AI development. Under current circumstances, any company with any CEO with any personnel would largely be incentivised to take the path that Anthropic, OpenAI etc. have taken.
The problem though with many (former) insiders who seek to spread this message is that they are often vague in their prescriptions and descriptions of what exactly the problem is, like Jacob Coxon. I understand they are worried about possibly commiting a crime by sharing company data to back their point, but if they truly believe in what they preach, one would think they would take the risk to avoid catastrophic situations.
Within one week, the firm Calif created a zero-click hacking tool that could compromise WeChat accounts of targeted people according to recent NYT reporting. No matter how well intentioned such an exercise is from the start-up, and Calif does indicate that its tools was built experimentally to benefit cyberdefences, it seems impossible to prevent such technologies from being weaponised in the future by states, and for other states to interpret it as such (everyone can imagine how a tool specifically attacking a WeChat vulnerability must look to the Chinese government..).
Anthropic's September report on misuse of its models poses interesting questions about AI safety. The report, detailing actions ranging from the use of Claude to set up a fake online dating profile farm to the model being employed by the Houthis to design software to guide missiles, was not produced under any legal obligation. Under the TFAIA in California and the EU AI Act, it is only mandated to confidentially disclose safety incidents to authorities, and the jury is still out on whether some of these cases would fall under the definition of 'serious incidents' (in the EU AI Act case) or 'critical safety incidents' (in the TFAIA case). Potential explanations for why the company still decided to produce such a report are not hard to imagine: better PR, internal employee pressure, and institutional inertia. However, these incentives are largely internal to the company and can change. If the board decides that these reports should be toned down to dissociate Anthropic from dangerous uses of AI (perhaps to prepare for an IPO), there is no reason why this would not take place. That would be a serious blow to the cause of AI safety at a time when the potential fallout from increasing model progress has still to be spelled out. Any loss of data on this front would mean one more threat vector left unexplored, and a part of society and the world left potentially unprepared.
I guess the discussion should not turn so much around shaming and blaming particular companies, but changing the structural incentives driving forward the current wave of AI development. Under current circumstances, any company with any CEO with any personnel would largely be incentivised to take the path that Anthropic, OpenAI etc. have taken.
The problem though with many (former) insiders who seek to spread this message is that they are often vague in their prescriptions and descriptions of what exactly the problem is, like Jacob Coxon. I understand they are worried about possibly commiting a crime by sharing company data to back their point, but if they truly believe in what they preach, one would think they would take the risk to avoid catastrophic situations.
Within one week, the firm Calif created a zero-click hacking tool that could compromise WeChat accounts of targeted people according to recent NYT reporting. No matter how well intentioned such an exercise is from the start-up, and Calif does indicate that its tools was built experimentally to benefit cyberdefences, it seems impossible to prevent such technologies from being weaponised in the future by states, and for other states to interpret it as such (everyone can imagine how a tool specifically attacking a WeChat vulnerability must look to the Chinese government..).