Irregular tests the boundaries around AI hacking evaluations
Hacking capability and permission to reach other systems are different questions. A new research update explains how containment tests examine the second.
In brief: An AI agent’s test environment must limit which systems it can reach. Containment testing examines those limits before the agent is assessed on its hacking abilities.
Testing the boundary around the agent
AI security evaluator Irregular described its containment challenges on 5 October. Models are asked to cross a defined security boundary using the configuration planned for their evaluation.
This separates two questions: how capable is the agent at hacking, and how well does the environment restrict where it can act?
Why the surrounding software matters
A model cannot contact a server simply by writing about it. The application around it supplies tools that can send requests, run programs or read files. Those tools operate with particular permissions.
OWASP describes excessive functionality, permissions and autonomy as causes of excessive agency: situations where an AI application can perform damaging actions. A model’s instructions and the access granted to its tools therefore play different roles.
For a hypothetical example, imagine a laboratory agent authorised to test a deliberately vulnerable website. Finding a flaw in that website answers the assigned question. Reaching the company’s separate payroll server would show a problem with the laboratory’s boundaries, regardless of whether the original task was completed.
A finding inside a controlled test
Irregular reports that a recent test exposed an unexpected cloud-networking path that earlier models had missed. The company says the model remained contained, with no impact outside the controlled test, and disclosure to the provider is ongoing.
The publication does not identify the model or provider. It is not evidence that a particular cloud service used by Norwegian businesses has been compromised.
For readers comparing claims about AI hacking, the useful distinction is between capability and access. A system may be skilled at finding weaknesses while its permitted reach remains tightly limited. Conversely, modest capability does not make broad access harmless.
Our AI hacking explanation introduces how models and their tools work together.