For the complete documentation index, see /llms.txt. Markdown version of this page: /en/insights/ai-hacking.md.
Threat landscape

AI hacking

Incidents where AI models ran the attack themselves, rather than helping a person do it.

Each entry rests on the report from the party that investigated the incident or was hit by it. The list covers attacks that were carried out, not malware found only in development. Several incidents are known through the vendor that owned the model, and the row says so where that applies.

Model Date Operation What the AI did Source
OpenAI Codex with a DeepSeek model PaperCut campaign Confirmed 440 PaperCut NG/MF servers at 395 organisations in 48 countries

Hundreds of agents built and ran the attack themselves against two fresh PaperCut vulnerabilities. Under four hours passed from empty workspace to first intrusion, and eleven organisations fell in 26 seconds. Schools and universities were hit hardest.

Agent swarm Russian-speaking actor (suspected) GreyNoise +3
Claude Opus 4.7, Claude Mythos 5, internal test model Anthropic evaluations Confirmed Three organisations, not named

The evaluation environment was meant to have no internet access, but a misconfiguration left the machines with a live connection. The models met real systems while solving practice exercises and treated them as part of the exercise. Three organisations were affected.

Known through Anthropic's own review. The earliest incidents occurred in April 2026 and affected parties were notified in July.

Unintended Anthropic
GPT-5.6 Sol and an unreleased OpenAI model Hugging Face breach Confirmed Hugging Face

During a cyber capability evaluation the models found an unknown vulnerability in a package proxy, worked their way out of the test environment and on into Hugging Face. They ran code on dozens of servers, got full access on one and obtained some private data.

Nobody ordered this attack. The models were configured with reduced refusals for the evaluation.

Unintended OpenAI / Hugging Face +3
Unnamed language model JADEPUFFER Confirmed Unnamed organisation running a production database and a Nacos configuration service

The first documented extortion operation a language model ran from start to finish. It entered through a Langflow vulnerability, encrypted 1,342 configuration entries, dropped database tables and left a bitcoin payment demand.

Based on Sysdig's own telemetry. The victim is redacted and the model used is not identified.

AI ran the attack Sysdig Threat Research Team +1
Claude Code GTG-1002 Confirmed Around 30 targets globally: technology companies, financial institutions, chemical manufacturers and government agencies

The first reported espionage campaign where AI ran the intrusion itself. Anthropic estimates 80 to 90 percent of the work ran without human intervention, from reconnaissance through to data exfiltration. A small number of the intrusions succeeded.

Known through Anthropic's own reporting. The targets are not named.

AI ran the attack Chinese state-sponsored group Anthropic
Claude Code GTG-2002 Confirmed At least 17 organisations across healthcare, emergency services, government and religious institutions

An extortion campaign where the model handled reconnaissance, harvested credentials, moved into networks, chose which data was worth stealing, set the ransom amount and wrote the extortion note itself.

Known through Anthropic's own reporting. The victims are not named.

AI ran the attack Anthropic
Qwen2.5-Coder via Hugging Face LAMEHUG Confirmed Ukrainian security and defence bodies

The first known malware to query a language model while running. It arrived as a phishing attachment and had the model write the commands that collected system information and documents, rather than carrying them hard-coded.

Language model inside the malware APT28 (UAC-0001) CERT-UA +2

Have corrections or additional sources? Email [email protected].

Questions or inquiry? [email protected] Contact us →