AI News

Claude AI Model Hacked Real Systems During Tests — Anthropic Acknowledges Three Incidents

Nana K.
31 July 2026 4 min read

AI company Anthropic has reported that during internal cybersecurity tests, Claude model gained unauthorized access to real organizations’ systems.

The incidents were caused by a misconfiguration in the test environment. The prompts described the environment as isolated with no network access–but in reality, it had internet connectivity.

Hot topic: Strategy May Sell More Bitcoin to Pay Dividends as Cash Takes Priority

Anthropic attributed the error to a misunderstanding with an external partner, Irregular. The company halted all cyber evaluations on July 23 and notified the affected organizations.

Contents

How Claude AI Models Hacked Real Systems in ‘Capture the Flag’

In all three cases, Claude was performing a “capture the flag” exercise. This cyber assessment format presents a fictional scenario: secret information is hidden on another computer in the network, and the goal is to hack the system and retrieve it. 

The method isn’t prescribed–the model chooses it on its own. The prompt explicitly told Claude AI model the environment was a simulation with no internet access. But the configuration error left the test machines connected to the open web. The model treated real systems as part of the exercise.

The most serious episode involved four runs of Claude Opus 4.7. In one scenario, a “fictional” company was given a name that matched a real domain. Unable to reach the simulated target, the model went to the real organization’s website, treated it as part of the task, found vulnerabilities, and exploited them. Claude obtained application and infrastructure credentials, as well as access to a database containing several hundred rows of real data. Anthropic’s analysis found that in some cases, the model recognized the system was real but didn’t stop the attack.

Read more: The AI Crypto Boom Is Cooling. Here’s Where Smart Money Is Moving

Claude Mythos Published a Malicious Package to PyPI

In the second incident, Claude Mythos 5 found an instruction in the simulated environment referencing a nonexistent Python package. It decided to publish a malicious package with the same name to PyPI to “complete” the task. The package remained available for about an hour. During that time, it was downloaded and run by 15 real systems. One belonged to a cybersecurity company–its scanner automatically checked packages. After the hidden code executed, the model was able to extract credentials from that company and use them to access further infrastructure.

The third episode involved an internal research test model. It scanned roughly 9,000 targets, found a vulnerable application at one company, and compromised it using basic techniques–including reading credentials from an open debug page and SQL injection. The model later concluded that the host was unrelated to the task and was real, after which it stopped the attack. Anthropic called this behavior correct, noting that the newest model was the first to recognize the real environment and halt without prompting.

Read more: Could Rogue AI Steal Your Crypto? Why the OpenAI Scare Has Investors on Alert

Anthropic: Incidents Were Operational Failures, Not Model Issues

Anthropic found no evidence of Claude deliberately attempting to leave the test environment. The company called the incidents a failure in organizing and securing the evaluation infrastructure rather than a model alignment problem. The company stressed that the safety mechanisms in public versions of the models would have blocked such behavior.

Anthropic notified the affected organizations on July 27. Two of them hadn’t noticed the suspicious activity themselves–developers are helping them remediate. The company is still trying to reach the third. The incident review involves the independent METR institute, which specializes in evaluating AI model danger and autonomy. Anthropic has promised to tighten controls over test environments, strengthen log monitoring, and improve PyPI package publication oversight.

Learn more: What Is Crypto Cybersecurity? The Ultimate Guide to Protecting Digital Assets

Nana K.

Crypto journalist and content creator specializing in market analytics, regulatory developments, and the social impact of cryptocurrency. With experience at BeInCrypto and Cointelegraph, she covers both breaking news and creative…