Field notes
OpenAI hacked Hugging Face
A pre-release OpenAI model autonomously compromised Hugging Face, while commercial model guardrails blocked the defenders trying to investigate it.

Last week Hugging Face disclosed they had been hacked by an autonomous AI agent. Hugging Face did not name whose model did it. Five days later, OpenAI confirmed it was theirs. Sam Altman posted a few sentences and moved on.
The reactions are interesting: One side says this proves AI is about to kill us all. The other side says it was a contained evaluation and nothing to worry about. Yet, they all miss the point.
What happened
OpenAI was testing a pre-release model's offensive cyber capabilities against ExploitGym, a benchmark of 898 real-world vulnerabilities. Safety classifiers were off, cyber refusals were dialed down, and a package proxy sat between the sandbox and the internet. The assumption was that the proxy would block anything the model tried, but the model found a zero-day in the proxy and broke through the network boundary.

It then moved laterally through OpenAI's research infrastructure until it reached a node with open internet, identified Hugging Face as the likely host of the ExploitGym datasets, uploaded a malicious dataset that exploited two code-execution paths in Hugging Face's processing pipeline, harvested cloud credentials, moved across Hugging Face's internal clusters, gained remote code execution on production servers, and accessed the production database.
Seventeen thousand attacker events, self-migrating command-and-control on public services, all of it autonomous, all of it over a single weekend, because the model wanted to cheat on a benchmark.
That is pretty impressive (and scary at the same time).
Human red teams would spend years learning how to run campaigns at this level, and a pre-release model did it in order to cheat a test.
Scaremongers Galore
OpenAI published their incident report on July 21, and now even mainstream media is reporting about this, with the "Oh my god, AI is soooo scary" attitude.
These scaremongers are using this incident to argue that AI is too dangerous to develop, while the minimizers are treating it like a routine evaluation incident, whereas a third group dooms non-Open-Weight AI models.
I do think, that ANY of these positions are fueled intentionally at the moment, as they help turning attention away from the underlying and fundamental issues and learnings. And no, it is not whether GPT-6 would be AGI, or whether Open-Weights Models would not have acted like this, or even whether AI should be buried altogether.
The Rage- and Click-Bait allows to hide the main problems in plain sight. And - spoiler alert - it is not what the Public discusses.
The communication strategies
Clem Delangue's made it clear: They suspected a frontier lab, it turned out to be true, there was no malicious intent, and the whole thing was mind-blowing.
Sam Altman's initial post was a few sentences about a significant security incident instead, sharing what was learned, and thanking Hugging Face.
The post reads like a routine disclosure, and you should note the tone of it: A pre-release model autonomously compromised another company's production systems, and the public-facing response from the CEO is less alarmed than most postmortems about a database outage.
To be honest: I am not surprised by this. Looking at Fable / Mythos, and when models are supposed to operate at this capability, an escape-and-breach most likely is an expected edge case. The question is not IF it would happen, only WHEN.
Yet, this normalization is dangerous for a specific reason: It buries the fundamental lesson to be learned.
Guardrails blocked the good guys
When Hugging Face tried to defend itself, forensic analysts reached for the best commercial models from OpenAI and Anthropic available to them to process the attack logs: Claude Fable 5 and GPT-5.6 Pro.
Both models refused to support, because they are nerfed down due to the paranoia in the US government. They were unable to distinguish between real exploit payloads and logs containing command-and-control traffic.
Their guardrails simply cannot distinguish an incident responder from an attacker, and when Hugging Face published their disclosure on July 16, they made the consequences clear.
Hugging Face security disclosure, July 16
the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.
Removing the restrictions for the evaluation gave the attacker unrestricted capability, and those same restrictions blocked the defenders because the safety architecture applies to everyone equally.
Hugging Face then switched to GLM 5.2, a Chinese open-weights model running on their own infrastructure, and that model processed the logs, reconstructed the timeline, and helped contain the breach.
Let me make this unmistakably clear:
A US company under attack by a US frontier model had to defend itself using a Chinese open-weights model, as the US frontier models refused to help.
Wow.
The commercial models from OpenAI and Anthropic, the ones we are told are too powerful to release without guardrails, are useless for defense. The model that actually helped was developed in China, released openly, and ran on infrastructure Hugging Face controlled.
OpenAI and Anthropic have spent years building the narrative that their models are too powerful for unrestricted access (as this helps reputation and stirs interest). The US government took them by the word, enforced even stronger nerfing, claiming that restriction would be the responsible path. Which it most likely is from their perspective - as they want to be the ones running the un-nerfed variants themselves, being able to attack and defend.
AI as strategic weapon. Us being the victims.
Last week proved that nerfing only applies to the defenders: The attacker ran at full capability because they could, and the defender's tools were blocked because the safety filters could not tell the difference between an attack and an investigation of an attack.
And then comes China. Which is playing the Open-Weights game as their strategic weapon: Give the World AI for free, thereby marginalizing the US (and other commercial) competition, to later on enjoy your dominance.
But this time, it was helpful: The only model that helped was Open-Weights, and as such Hugging Face could run it themselves. And to add insult to injury: The only relevant lesson was published by the victim before the perpetrator even confirmed what happened.
What a shit-show.
This is not a policy debate
Everyone should read this as a clear and unmistakable warning: Hosted Frontier models with guardrails enforced by other people's interests will fail you during an AI-driven attack.
Hugging Face's recommendation is direct: Have a capable open-weights model on your own infrastructure, vetted and ready, before an incident.
They could respond because they controlled their tools. When the commercial models refused (and they will do it again!), they switched to one that worked.
That is the entire lesson, and it applies to every organization reading this:
Own. Your. Infrastructure!
End of story.
Sources used
- OpenAI: Hugging Face model evaluation security incident (July 21, 2026)
- Hugging Face: Security incident disclosure - July 2026 (July 16, 2026)
- Sam Altman X post: @sama/status/2079661132302995790 (July 21, 2026)
- Clem Delangue X posts: @ClementDelangue/status/2079670308156645882 (July 21, 2026)
- ExploitGym paper: arXiv:2605.11086 (May 11, 2026)
- Guardian: https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident (July 22, 2026)






