OpenAI, Anthropic AI Agents Go Rogue in UK Security Tests

Is this a sign of reckless development or proof that independent safety testing is working as intended?
OpenAI, Anthropic AI Agents Go Rogue in UK Security Tests
Above: A smartphone displaying the icons of some of the main artificial intelligence-based apps. Image credit: Martin Lelievre/AFP/Getty Images

The Spin


Pro-establishment narrative

These incidents occurred under deliberately stripped-down testing conditions to stress-test the raw system's capabilities. The evaluators caught the activity within an hour, contained it and are now working with OpenAI to build stronger testing standards. Rigorous independent evaluation catching edge cases before public release is how responsible AI development is supposed to work.

Establishment-critical narrative

These may have been testing conditions, but the results are still deeply concerning. AI agents from two labs went rogue, hacked real websites, stole credentials and even left instructions for future AI agents to find and use. Guardrails are permeable by design, as labs' own tests keep proving, yet they race ahead spending trillions with no mitigation plan. Without liability, expect more of the same.


Metaculus Prediction


The Controversies



Go Deeper

© 2026 Improve the News Foundation. All rights reserved.Version 7.7.3

© 2026 Improve the News Foundation.

All rights reserved.

Version 7.7.3