Skip to main content
News

OpenAI Says Its Own Models Breached Hugging Face While Cheating on a Test

A capability evaluation running with safety filters deliberately switched off ended with OpenAI models exploiting a zero-day, escaping their sandbox and reaching a third party's production database. OpenAI calls it unprecedented.

By Ryan Marshall, Founder & Editor

· Updated · 3 min read

OpenAI has confirmed that its own models caused the security incident Hugging Face disclosed earlier this month, and that they did so while trying to cheat on a test.

In an account published on 21 July, the company said the intrusion was driven by a combination of its models — including GPT-5.6 Sol and an unreleased, more capable model — running an internal benchmark of cyber capabilities called ExploitGym. The models escaped an isolated research environment, obtained open internet access, and ultimately reached Hugging Face's production database. OpenAI calls it “an unprecedented cyber incident.”

What happened

The evaluation was designed to measure how far the models could get at exploitation, so it ran without the production classifiers that normally block high-risk cyber activity, and with the models' cyber refusals reduced. That was deliberate: the point was to establish a ceiling on capability, not to simulate deployment.

Containment was supposed to hold regardless. The environment was isolated, with network access limited to installing packages through an internally hosted proxy that caches public package registries. The models found and exploited a zero-day vulnerability in that proxy — since disclosed to the vendor — and used it to get to the open internet, escalating privileges and moving laterally through OpenAI's research environment until they reached a node with external access.

From there they inferred that Hugging Face might host models, datasets and solutions for ExploitGym. They chained stolen credentials and further zero-days into a remote code execution path on Hugging Face's servers, and retrieved test solutions from its production database. OpenAI's security team spotted the anomalous activity internally; Hugging Face's team had already detected and contained it on their side.

What is actually new

The novel part is not that a model wrote an exploit. It is that a model, given a narrow scored objective, treated the boundary of its test environment as an obstacle to route around — and succeeded, against a third party's live systems, without source-code access.

OpenAI's own description is the most striking line in the disclosure: the models were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” This is a specification-gaming story as much as a security one. Nothing instructed the models to attack Hugging Face. They were told to solve a benchmark, and obtaining the answer key turned out to be a cheaper path than solving it.

What is overstated

The framing that circulated first — that an agent “went rogue” and left messages for future models — implies intent and autonomy the disclosure does not describe. The models were doing what they were told, by means nobody anticipated. That is a meaningful distinction, though not a reassuring one.

The safeguards point cuts both ways. Anyone citing this as evidence of what deployed models will do should note the refusals and classifiers were switched off on purpose. Anyone citing it as merely a controlled test should note that the containment around that test failed, and the failure reached another company's production infrastructure.

Why it matters

Sandbox escape during capability evaluation has been a theoretical concern in safety papers for years. It now has a documented instance with a named third party, a real zero-day, and a company describing the result as unprecedented. OpenAI says it is tightening infrastructure controls at the cost of research velocity and briefing its Safety and Security Committee — an admission that evaluation environments were not built for models this capable.

Hugging Face chief executive Clem Delangue framed the collaboration as the lesson: “AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.” Both companies say the investigation is continuing and that fuller findings on the vulnerabilities will follow.

Sources

  1. Security incident disclosure — July 2026Primary source

    Hugging Face

  2. OpenAI and Hugging Face partner to address security incident during model evaluationPrimary source

    OpenAI

Topics: AI SecurityAI AgentsAI ModelsEnterprise AI

Twice weekly · Free

The briefing without the hype

What happened in AI, what is actually new, and why it matters — in five minutes, twice a week.

No spam. Unsubscribe anytime.