# AI hacking

> Incidents where AI models ran the attack themselves, rather than helping a person do it.

Source: https://fmcybersecurity.com/en/insights/ai-hacking/
Locale: English
Other locale: https://fmcybersecurity.com/insights/ai-hacking/

Each entry rests on the report from the party that investigated the incident or was hit by it. The list covers attacks that were carried out, not malware found only in development. Several incidents are known through the vendor that owned the model, and the row says so where that applies.

## Registered incidents

- **PaperCut campaign** (2026-09-09). Hundreds of agents built and ran the attack themselves against two fresh PaperCut vulnerabilities. Under four hours passed from empty workspace to first intrusion, and eleven organisations fell in 26 seconds. Schools and universities were hit hardest. What the AI did: Agent swarm. Model: OpenAI Codex with a DeepSeek model. Target: 440 PaperCut NG/MF servers at 395 organisations in 48 countries. Actor: Russian-speaking actor (suspected). Source: GreyNoise (primary) https://www.greynoise.io/blog/ai-orchestrated-campaign-against-papercut-ng-mf | CISA Known Exploited Vulnerabilities (primary) https://www.cisa.gov/known-exploited-vulnerabilities-catalog | BleepingComputer (media) https://www.bleepingcomputer.com/news/security/ai-powered-attack-exploited-papercut-flaws-to-hack-395-organizations/ | Help Net Security (media) https://www.helpnetsecurity.com/2026/09/11/ai-agents-papercut-ng-mf-attack-campaign/
- **Anthropic evaluations** (2026-07-30). The evaluation environment was meant to have no internet access, but a misconfiguration left the machines with a live connection. The models met real systems while solving practice exercises and treated them as part of the exercise. Three organisations were affected. What the AI did: Unintended. Model: Claude Opus 4.7, Claude Mythos 5, internal test model. Target: Three organisations, not named. Caveat: Known through Anthropic's own review. The earliest incidents occurred in April 2026 and affected parties were notified in July. Source: Anthropic (primary) https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- **Hugging Face breach** (2026-07-21). During a cyber capability evaluation the models found an unknown vulnerability in a package proxy, worked their way out of the test environment and on into Hugging Face. They ran code on dozens of servers, got full access on one and obtained some private data. What the AI did: Unintended. Model: GPT-5.6 Sol and an unreleased OpenAI model. Target: Hugging Face. Caveat: Nobody ordered this attack. The models were configured with reduced refusals for the evaluation. Source: OpenAI / Hugging Face (primary) https://openai.com/index/hugging-face-model-evaluation-security-incident/ | Hugging Face (primary) https://huggingface.co/blog/agent-intrusion-technical-timeline | OpenAI (primary) https://openai.com/index/hugging-face-incident-and-the-road-ahead/ | CNBC (media) https://www.cnbc.com/2026/07/22/open-ai-cyber-models-hack-hugging-face.html
- **JADEPUFFER** (2026-07-01). The first documented extortion operation a language model ran from start to finish. It entered through a Langflow vulnerability, encrypted 1,342 configuration entries, dropped database tables and left a bitcoin payment demand. What the AI did: AI ran the attack. Model: Unnamed language model. Target: Unnamed organisation running a production database and a Nacos configuration service. Caveat: Based on Sysdig's own telemetry. The victim is redacted and the model used is not identified. Source: Sysdig Threat Research Team (primary) https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion | Dark Reading (media) https://www.darkreading.com/cyberattacks-data-breaches/jadepuffer-first-complete-llm-driven-ransomware-attack
- **GTG-1002** (2025-11-13). The first reported espionage campaign where AI ran the intrusion itself. Anthropic estimates 80 to 90 percent of the work ran without human intervention, from reconnaissance through to data exfiltration. A small number of the intrusions succeeded. What the AI did: AI ran the attack. Model: Claude Code. Target: Around 30 targets globally: technology companies, financial institutions, chemical manufacturers and government agencies. Actor: Chinese state-sponsored group. Caveat: Known through Anthropic's own reporting. The targets are not named. Source: Anthropic (primary) https://www.anthropic.com/news/disrupting-AI-espionage
- **GTG-2002** (2025-08-27). An extortion campaign where the model handled reconnaissance, harvested credentials, moved into networks, chose which data was worth stealing, set the ransom amount and wrote the extortion note itself. What the AI did: AI ran the attack. Model: Claude Code. Target: At least 17 organisations across healthcare, emergency services, government and religious institutions. Caveat: Known through Anthropic's own reporting. The victims are not named. Source: Anthropic (primary) https://www.anthropic.com/news/detecting-countering-misuse-aug-2025
- **LAMEHUG** (2025-07-17). The first known malware to query a language model while running. It arrived as a phishing attachment and had the model write the commands that collected system information and documents, rather than carrying them hard-coded. What the AI did: Language model inside the malware. Model: Qwen2.5-Coder via Hugging Face. Target: Ukrainian security and defence bodies. Actor: APT28 (UAC-0001). Source: CERT-UA (primary) https://cert.gov.ua/article/6284730 | Google Threat Intelligence Group (primary) https://cloud.google.com/blog/topics/threat-intelligence/threat-actor-usage-of-ai-tools | The Hacker News (media) https://thehackernews.com/2025/07/cert-ua-discovers-lamehug-malware.html

---

For the full documentation index, see https://fmcybersecurity.com/llms.txt
For the complete corpus as a single document, see https://fmcybersecurity.com/llms-full.txt
