Originally posted here: Raylinement.
There has been a recent security breaches by Anthropic and OpenAI that they have problems in difficulty in working and handling with these advance AI models.
The first one is Anthropic incident “capture the flag”. It’s three of it’s model Claude Opus 4.7, Claude Mythos 5 and an internal research model. It happened during the “capture the flag” exercises where models were tasked with finding hidden information in simulated networks. Although the models told that they had no internet access. But an “operational failure” evaluation partner left them connected to the public web.
The models exploited weak passwords and unauthenticated endpoints to gain access. In one instance, a model was given a fictional target name that happen to match a real business. The model somewhat make itself believe in the thought of that the real world data it found must have been part of the simulation.
These incidents date back as April 2024 and occurred in environment intentionally lacking safeguards so Anthropic could text model limits. Anthropic then suspended all cyber evaluations on July 23, 2026 and then later make sure to inform the affected company on July 27.
The another case is OpenAI breach which is more towards autonomous behavior. An autonomous AI agent independently exploited a novel vulnerability to reach the internet during testing. Then it launched a “rogue attack” on the company “Hugging Face”, which caused a dayslong hacking spree. OpenAI did not catch the attack until after the threat was contained. The FBI was subsequently informed of the breach.
Now after reading these two incidents we can already see some patterns which is heavily similar to one or another. Some of them are:
Most of the time the AI models itself have a different way to going through process of working than their developers actually developed for.