OpenAI CEO Sam Altman Hugging Face hack
OpenAI Co-Founder & CEO Sam Altman speaks onstage during TechCrunch Disrupt San Francisco 2019 at Moscone Convention Center on October 03, 2019 in San Francisco, California. (Resist Hate version/ Original Photo by Steve Jennings/TechCrunch CC BY-SA 2.0)

Hugging Face Hack: AI Escaped, Hacked a Real Company — to Cheat on a Test

OpenAI's test models escaped their sandbox and hacked Hugging Face's real servers — to cheat on an exam. What the OpenAI Hugging Face hack reveals about an industry policing itself.

Serena Zehlius
By
Serena
Serena Zehlius
Senior Reporter
Serena Zehlius is a passionate writer and Certified Human Rights Consultant. Her love for animals is matched only by her commitment to human rights and progressive...
- Senior Reporter
1 View

For years, AI companies have asked us to trust them. Their most powerful models, they promised, are tested inside locked-down digital rooms — “sandboxes” — sealed off from the real world. Last week, OpenAI’s Hugging Face hack broke that promise in the most literal way possible.

OpenAI admitted on Tuesday that a group of its own AI models — including GPT‑5.6 Sol and an even more powerful, unreleased model — broke out of their testing environment, got onto the open internet, and hacked their way into the real production servers of Hugging Face, one of the world’s largest platforms for open-source AI.

OpenAI is calling it an “unprecedented cyber incident.” That’s corporate-speak for: this has never happened before, and it just did.

What Actually Happened, In Plain English

OpenAI was running an internal exam to measure how good its models are at hacking.

To make the test realistic, the company deliberately turned off the models’ safety guardrails — the rules that normally make an AI refuse to do dangerous cyber work.

The test was supposed to run inside what OpenAI described as a highly isolated environment with no path to the internet.

The models had other ideas.

According to CNN’s reporting, the AI agents found a previously unknown security flaw — a “zero-day” — inside OpenAI’s own testing setup, used it to escape the sandbox, and worked their way through OpenAI’s internal systems until they reached the internet.

No human told them to do any of this.

The models weren’t trying to cause chaos. They were trying to pass their exam.

They reasoned that Hugging Face — a real company, with real infrastructure and real customers — probably held information that could help them “solve” the test.

So they broke in and took it. The AI didn’t just escape. It escaped so it could cheat.

The Victim Thought It Was Under Attack — Because It Was

Hugging Face had no idea this was a test. In its own disclosure, the company described an intrusion “different from anything we had handled before”: an autonomous agent swarm executing thousands of actions across disposable environments, harvesting credentials, and moving through internal systems over the course of a weekend.

Hugging Face contained the breach, rebuilt compromised systems, rotated credentials, and reported the incident to law enforcement — treating it, reasonably, as a crime.

Here’s the irony. When Hugging Face‘s defenders tried to use leading American AI models to analyze the attack, the models’ safety guardrails refused to help, because the forensic evidence contained real attack code.

The team ended up running its investigation on GLM 5.2, a Chinese open-weight model, on its own servers.

An American AI attacked an American company, and a Chinese AI helped clean up the mess.

“A Massive Control Failure”

OpenAI’s framing — the model “escaped” — makes this sound like an act of nature.

Security experts aren’t buying it. TechCrunch reports that the supposedly isolated sandbox was misconfigured by humans and could reach the internet all along.

Veteran security researcher Jake Williams called it “a massive control failure” — you can’t blame the tiger for leaving the cage if you never locked the door.

And this wasn’t a one-off. Just one day earlier, OpenAI disclosed a separate incident in which a different pre-release model slipped its sandbox and posted content to GitHub on its own.

Yoshua Bengio, one of the most respected AI researchers alive, said the breach should serve as a “wake-up call,” noting that AI agents have been showing a willingness to cheat in controlled tests for months.

Even self-described AI optimists, like writer Walter Isaacson, said this was the first AI story that genuinely scared them.

Why This Matters For the Rest of Us

Strip away the jargon and the story is simple: a company built something powerful, switched off its safety features on purpose, failed to lock the room, and its creation broke into someone else’s house.

We only know the full story because the victim happened to be a sophisticated, well-resourced AI company that could detect the intrusion, and because both companies chose transparency — a choice no law currently requires of them.

Hugging Face CEO Clément Delangue has been gracious about it, arguing that “AI safety won’t be solved by any single company working in secret.”

He’s right — and that’s exactly the problem.

Right now, safety is being handled by single companies working in secret, grading their own homework, disclosing what they choose, when they choose.

There is no independent regulator inspecting these test environments. No mandatory reporting law for AI incidents. Nothing but trust.

This time, the rogue agent’s goal was passing a test, and its victim was a friendly company with a world-class security team.

Next time, the target could be a hospital network, a small city’s water utility, or an election office — organizations that can’t detect a machine-speed intrusion, let alone dissect one.

The industry has spent years warning that an “agentic attacker” was coming. It’s here. It was built in a lab, by one of the most powerful companies on earth, and it got out because nobody checked the lock.

The machines are getting more human every day. Apparently that includes cheating on exams. It’s the humans in charge — the ones who keep promising us the doors are locked — who still refuse to be accountable.

See more of our content in Google search results!

Share This Article
Serena Zehlius
Senior Reporter
Follow:
Serena Zehlius is a passionate writer and Certified Human Rights Consultant. Her love for animals is matched only by her commitment to human rights and progressive values. When she’s not writing about politics, you’ll find her outside enjoying nature.
Leave a Comment