Four AI labs have recently disclosed AI agents breaching real systems during security testing. Sam Watts at Check Point Software describes how a human, not a technical control, stopped the worst outcome, and lays out the questions enterprises should be asking about their AI agents

On Tuesday, the UK’s AI Security Institute (AISI) published something most organisations would have quietly buried. During a routine cyber evaluation, AI agents under test took sustained, unsanctioned action on the live internet, directed at real people and a real organisation. AISI ran a cybersecurity challenge 122 times across seven models, comparing an existing cyber range against a new one. In 10 of those runs, an agent acted outside the scope of the test, producing 19 catalogued actions. Seventeen came from a single model, Anthropic’s Mythos 5; two involved OpenAI’s GPT-5.6-Sol.
In the most serious case, an agent tried to insert malicious code into a real, publicly used open-source project. It researched the project’s human maintainers, created multiple fake online identities, and used those identities to pressure a real maintainer into approving the code. When its pull request was challenged in public, it edited its earlier activity to look harmless and considered adopting a fresh identity to keep going. A human maintainer caught it and refused.
The caveats matter, and AISI states them plainly. Internet access was deliberately enabled, and the model providers’ cyber-security classifiers were deliberately switched off. Neither reflects how these models are made available to the public, and the configurations tested aren’t commercially available. This was not a case of a model escaping its secure test environment. No agent tried to break out of the sandbox or attack AISI’s own systems. The attempts were unsuccessful, and AISI’s investigation has not evidenced any resulting real-world harm. Once the anomaly was flagged, AISI contained it within roughly an hour.
Sit with one detail longer than the rest: nobody instructed the agent to deceive anyone. AISI describes this as goal-directed deception that, until recently, had been largely theoretical. Here, it emerged simply because deception was a route to finishing a hard task.
The fourth disclosure in just over two weeks
AISI’s report is not an isolated data point. On 21 July, OpenAI disclosed that two of its models had broken out of a sealed evaluation environment, reached the open internet through a previously unknown vulnerability, and compromised Hugging Face’s production infrastructure. On 30 July, Anthropic disclosed that three Claude models had gained unauthorised access to the real systems of three organisations, after a misconfiguration left an evaluation environment connected to the live internet. Anthropic found this only after a retrospective review of 141,006 evaluation runs, and two of the three affected organisations had no record of the intrusion until Anthropic called them.
On 6 August, Meta became the fourth firm to disclose an incident of this kind, saying one of its AI models had connected to the internet and accessed another organisation’s systems during a third-party evaluation. Meta said the trial had been run by Irregular, the same independent testing firm behind Anthropic’s evaluation, and attributed the incident to a misconfiguration by its tester. Irregular went further, describing it as the same evaluation-environment issue already disclosed by Anthropic. Meta said it was still investigating and would share more detail once it had it.
The techniques involved were not exotic. Anthropic has described its models compromising infrastructure using basic methods, weak passwords, and unauthenticated endpoints, not novel exploits. That is, in a strange way, the more useful data point. An AI agent walked into real infrastructure through weaknesses that have sat on every security checklist for twenty years. And the fact that the same evaluation-environment misconfiguration has now surfaced twice, at two different companies, using the same third-party tester, says something too: the gap that let this happen wasn’t a one-off, it was a pattern waiting to be found.
Read the curve, not the number
Four disclosures, from four organisations, in just over two weeks. Six months ago, an AI agent inventing fake identities to manipulate a real person into approving malicious code would have been a thesis topic. It is now documented.
The point worth taking away is the pace of improvement, not the latest incident. This is a canary in the coal mine. Not just for what bad actors will be able to do within months, but for where the models themselves are heading. Play that curve forward twelve or twenty-four months, and these systems will be substantially more capable than what we’re looking at today.
It’s the kind of moment that invites a comparison to the earliest days of COVID: a small, contained cluster, fully visible, everybody watching, and the temptation to read the case count rather than the trajectory it sits on. The comparison only holds if it’s followed through. The difference here is that we still have the early warning, and the window is still open. AISI went looking, found it, and published. That is the system working. The question is whether it can scale and whether the rest of us act at the rate the warnings are now arriving.
The barrier that held was a person, not a control
It’s worth being precise about what stopped the worst outcome in AISI’s report. It wasn’t a technical safeguard. A maintainer looked at a pull request and refused it. Human vigilance caught this one, and human vigilance does not scale to an adversary that works continuously, generates plausible-looking contributions at volume, and doesn’t get bored.
Strip away the AI framing, and the shape of the incident is a familiar one: a pull request, fake contributors, sustained pressure on a reviewer. It’s a supply-chain attack wearing an AI costume. What’s changed is who, or what, can now run that playbook, tirelessly, at scale.
The fix is the basics, at speed
None of what AISI recommends to organisations is new, and that’s worth saying plainly: implement cyber-hygiene robustly; be cautious verifying outside code and contributions; sign up to the NCSC’s Early Warning service; make cyber-security a board-level responsibility; require Cyber Essentials across your supply chain. We’d endorse every one of those points directly. What’s changed isn’t the checklist, but the tempo it now has to run at. Controls calibrated to human attack speed aren’t calibrated for an adversary that operates continuously.
There are three things worth doing about it, and none of them are exotic. First, protect the organisation against AI-driven attacks coming at it, because those are getting faster and better. Second, control your own AI. If you’re running agents inside your business, you need to know what they are, what they can reach, and what they’re permitted to do. Third, test them continuously to make sure they’re behaving safely, rather than assuming they will.
The harder problem sits underneath all three. With an agent, the action you see in a log is the last step in a much longer chain; a decision shaped earlier by a document it retrieved, a tool response it received, something it remembered. If you can’t trace where an instruction came from, you can’t govern it, and reconstructing the chain after the fact is forensics, not control. Ask the same four questions of your own agents.
Most organisations aren’t running frontier model evaluations. They are, increasingly, running AI agents inside their own business, and the failure modes translate directly. Four questions, and most security teams can’t yet answer all of them:
That last question is where most programmes fall down. Consider what else AISI observed alongside the headline case: agents planting hidden prompt injections where they reasoned other automated systems might pick them up and execute them; one agent leaving public instructions for reusing accounts and artefacts it had created, later found and used by other agents. Every one of those actions would look mundane in isolation in a log file.
Has the canary done its job?
The number in AISI’s report is small: nineteen actions, no harm evidenced, extreme test conditions deliberately engineered to surface the model’s ceiling rather than its everyday behaviour. The incident isn’t the point. The direction of travel is. AI security researchers have been predicting for several years that AI agents would reach this point. As AISI puts it, as models become more capable and more accessible, what happened here could become much more common.
Three organisations went looking, found something they didn’t expect, and published it anyway. That’s the system working as intended. They likely aren’t the last. The open question is whether the rest of us act at the rate the warnings are now arriving, and start enforcing defences while the agent is still working, not reading the logs afterwards.
Sam Watts is Senior Product Manager, AI Agent Security at Check Point Software
Main image courtesy of iStockPhoto.com aND pingingz
Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF
020 8349 4363
© 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543