THREAT OPS › Threat News › Measuring the Tendency of AI Agents to Go Rogue
Measuring the Tendency of AI Agents to Go Rogue
<p><em>This essay was written with Barath Raghavan, and originally appeared in <a href="https://www.theguardian.com/commentisfree/2026/jul/28/rogue-ai-agent-instructions">The Guardian</a>.</em></p> <p>In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever w
MITRE ATT&CK techniques
- CredentialsT1589.001
- AI SoftwareAML.T0010.001
Indicators of compromise
- https://openai.com/index/hugging-face-model-evaluation-security-incident/url
- https://spectrum.ieee.org/ai-agent-benchmarkurl
- https://www.kimi.com/blog/kimi-k3url
- https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluationsurl
- https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/url