For the complete documentation index, see /llms.txt. Markdown version of this page: /en/insights/ai-security/crowdstrike-ai-safety-checks-miss-combined-attacks.md.
AI Security ↗

CrowdStrike: Safe AI answers can become parts of an attack

A safety decision about one question does not necessarily account for the workflow its answer becomes part of.

AI-generated illustration of a laptop with the CrowdStrike logo and a diagram connecting three separate requests to a combined result.
AI-generated research illustration. The screen is a concept, not a CrowdStrike product screenshot.

In brief: An AI safety check can approve individual questions without seeing the harmful result their answers form together. CrowdStrike’s new research examines that gap.

What CrowdStrike found

In research published on 6 October, CrowdStrike reports that a tested safety classifier resisted direct bypass attempts. However, separately permitted assistance could be combined into offensive code across nine of ten tested security categories.

That is a result within the researchers’ evaluation, not a success rate for attacks against businesses.

Why an approved answer can still contribute to harm

A safety classifier is a screening system that assesses whether a request is allowed. Its decision depends on the information available to it.

Microsoft researchers describe the related problem as capability laundering. A system retaining the overall harmful goal can obtain help with ordinary subtasks from another model. That second model sees the separate questions, while the system combining the answers retains the reason they were asked.

The important distinction is between judging a piece of assistance and judging the activity it supports.

As a simple analogy, three people may each receive a routine work request. None sees the full project. Assessing each person’s answer cannot automatically establish whether the project itself is legitimate.

What this means for claims about AI safety

For a business comparing AI services, refusal of a harmful request is one useful result. It does not answer every question about a workflow involving several models, tools and steps.

Microsoft’s cybersecurity experiments used isolated benchmark environments. Neither that setup nor CrowdStrike’s published evaluation establishes how often this happens in real attacks. The findings concern the limits of what individual safety decisions can demonstrate.

CrowdStrike’s research is also separate from proof that a particular security product detects this behaviour. Our Falcon Guardian explanation covers the product’s stated role in protecting AI activity.

← Back to all insights
Questions or inquiry? [email protected] Contact us →