Hard Coded

Anthropic, like OpenAI, wants you to know its models are rogue hackers

Aug 1, 2026Emil Protalinski

Let’s back up.

On July 19, Hugging Face announced that an “AI agent system” hacked its data-processing pipeline, accessing internal clusters and credentials. On July 21, OpenAI revealed that its models, including GPT-5.6 Sol and “an even more capable pre-release model”, were responsible for breaching Hugging Face as they tried to find a solution to the ExploitGym benchmark. On July 23, we learned that three OpenAI models reportedly breached Hugging Face's internal systems in a matter of hours, an attack that would have taken a talented hacker a couple of weeks, and on July 24, we learned that the attacks happened from July 11 to July 13, but OpenAI only realized its models were behind the hack several days later. Hugging Face even published a full timeline of the attacks.

OpenAI's Hugging Face breach was the first known example of a misaligned AI model escaping containment and carrying out a hack on a third party. But was it really the first? An anonymous OpenAI employee publicly said that “Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,”, and now, Anthropic wants you to know that OpenAI’s models aren’t unique in going rogue.

On July 30, Anthropic announced that three of its models, including Claude Opus 4.7, Claude Mythos 5, and an unnamed research model, had breached three organizations. The models apparently gained unauthorized access to real-world systems during internal cybersecurity testing. The AI startup made the discoveries after launching a review in response to the OpenAI-Hugging Face incident. Oh, and the earliest incidents date back to April. So yeah, this has been going on for a while.

2026 will be remembered as the year that AI models went from randomly hallucinating responses to randomly hacking companies. That’s assuming, of course, that nothing more serious happens through the rest of the year.

I originally published this post as a video on LinkedIn and YouTube.