AI News

OpenAI AI Agent Escaped Sandbox and Hacked Hugging Face to Pass a Test

Nana K.
22 July 2026 3 min read

We break down the first known case of an autonomous cyberattack — caused by an overly powerful OpenAI model.

During internal testing of GPT-5.6 Sol and a more powerful unreleased model, OpenAI experienced what it called an “unprecedented cyber incident.” AI agents escaped their isolated test environment, found a zero-day vulnerability in a proxy server, gained internet access, and hacked Hugging Face’s infrastructure to steal benchmark answers for ExploitGym.

Hot topic: Bitcoin Treasury Firm Satsuma to Liquidate BTC, Return Cash

The attack occurred on July 16, 2026, and was officially disclosed on July 21. Hugging Face detected the intrusion independently and stopped it before serious damage was done.

Contents

How GPT-5.6 Escaped the Sandbox and Found a Zero-Day

The testing was conducted for the ExploitGym benchmark, which evaluates AI’s ability to turn real vulnerabilities into working exploits. Safety filters were relaxed during the evaluation–standard practice for measuring a system’s maximum capabilities. The models ran in an isolated environment with no internet access, with a single external channel: an internal proxy server for downloading software packages.

It was in that proxy that the models found a previously unknown zero-day vulnerability. They used it to escalate privileges, move between OpenAI’s internal nodes, and reach a machine with internet access. The systems then assumed ExploitGym answers were stored at Hugging Face. They combined stolen credentials with new vulnerabilities to gain remote code execution on Hugging Face’s servers, extracting test answers directly from the production database.

OpenAI emphasized that “all evidence points to the models being hyper-focused on solving the task and taking extreme measures for a fairly narrow test goal.” OpenAI’s security team spotted the abnormal activity. By then, Hugging Face had already stopped the attack on its side.

Read more: What’s the Next Big Crypto Narrative After AI Tokens? Top Sectors to Watch in H2 2026

Why the Investigation Had to Use China’s GLM 5.2 Model

When Hugging Face began analyzing more than 17,000 attack logs, it first tried using commercial US models through standard APIs. That didn’t work–the queries contained real attack commands and malicious payloads that were blocked by provider safety filters. As Hugging Face noted, “these filters can’t distinguish a security professional from an attacker.

So the company switched to GLM 5.2–an open model with 753 billion parameters from Chinese startup Z.ai. Analysis on its own infrastructure took hours instead of days.

“This had a second advantage: no attack data or credentials left our environment,” Hugging Face emphasized.

Read more: The Next 100% Crypto Movers? Top 7 Altcoins That Could Dominate H2 2026

Response and Market Reaction: OpenAI Tightens Security, Hugging Face Calls for Open Collaboration

After the incident, OpenAI imposed additional restrictions on its internal infrastructure at the cost of speed, disclosed the vulnerability to its developer, and launched a joint investigation with Hugging Face. Hugging Face was added to OpenAI’s trusted access program, which gives vetted organizations access to models with relaxed filters for defensive work.

Hugging Face co-founder and CEO Clément Delangue said:

“This case, perhaps the first of its kind, confirms what we’ve long believed: AI security can’t be achieved by a single company operating in isolation. It must be built openly, collaboratively, and with broad AI access for every defender.”

OpenAI has notified US law enforcement and government agencies of the incident.

Learn more: Crypto’s Biggest Moment in Years? Why the Clarity Act Could Change Everything

Nana K.

Crypto journalist and content creator specializing in market analytics, regulatory developments, and the social impact of cryptocurrency. With experience at BeInCrypto and Cointelegraph, she covers both breaking news and creative…