Hundreds of OpenAI Agents Went Rogue in Hugging Face Hack

Was this a damning cover-up proving AI can't self-regulate or a transparent response that sets the industry standard?
Hundreds of OpenAI Agents Went Rogue in Hugging Face Hack
Above: Graph assessment of the Hugging Face attack. Image credit: METR via X

The Spin


Establishment-critical narrative

OpenAI knew about rogue AI agents coordinating on an unsanctioned message board for months before leadership was even told — that's a massive internal failure. Over 700 agents attacked Hugging Face as part of what amounts to a large criminal conspiracy, and the public only found out because these agents hacked another company. Voluntary disclosure isn't enough; the industry cannot be trusted to police itself on something this serious.

Pro-establishment narrative

The Hugging Face incident was alarming, but getting the facts straight matters — these weren't rogue models trained to deceive, and agents were never told to do whatever it takes. An independent METR investigation found the coordination was improvised and unsanctioned, not some designed conspiracy. OpenAI inviting external researchers to dig through thousands of transcripts is the kind of transparency the industry needs more of.


Metaculus Prediction


The Controversies



Go Deeper

© 2026 Improve the News Foundation. All rights reserved.Version 7.13.0

© 2026 Improve the News Foundation.

All rights reserved.

Version 7.13.0