Many thanks to Ben Recht, Gautam Dasarathy, Narasimha Chari, Vi Iyengar for their review.
The OpenAI x Hugging Face incident was an extraordinary demonstration of security threats posed by cyber-enhanced frontier model powered agents. We are told that while solving an impossibly hard task, the model decided to cheat and orchestrated a dizzyingly complex attack on Hugging Face to steal the answer key. It is only normal to expect that a powerful model goes rogue to achieve its ends, means be damned.
Troubling Questions: It is hard not to anthropomorphize as you ponder the details. Why did the model decide to cheat instead of trying harder or declare impossibility? How did the idea of collaborating emerge? What intent could have propelled them to write coded messages in an illegally hacked message board? Are these agents in fact conscious if they recruited each other, engaged in “sacrifice” for the collective?
As more details followed, there were gaps that raised suspicion that something was amiss. There is a far simpler explanation, without attributing all to reward hacking.
Occam’s Razor
A purpose-built cyber-enhanced model trained on all the capabilities you saw (coordination, privilege escalation, cluster take-over, coordination, deception, stealth) was publicly demonstrated on a possibly carefully chosen target or was conveniently used to advance policy that favors the frontier labs.
Engineered, not “Emergent”
Models don’t “want” to hack. The evidence points to deliberate measurement, optimization and engineering of the model and the architecture around it to enable offensive capabilities demonstrated. OpenAI admits this is not meant for public release as they are designed for limited release for Cyber-defense.
Anthropic agrees too:
As LLMs scale in size, “emergent abilities” — skills that were not evident in smaller models and were not necessarily an explicit target of model training—appear. Indeed, Claude’s abilities to execute cybersecurity tasks like finding and exploiting software vulnerabilities in Capture-the-Flag (CTF) challenges have been byproducts of developing generally useful AI assistants.
Offensive Cyber-capabilities
Mere improvement in model capability hadn’t resulted in improvements in offensive cyber-capability. UK’s AI safety institute showed for instance that scaffolds matter. Google and Anthropic developed and touted their cyber-toolkits capable of offensive cyber-attacks and defense.
Training to Exploit and Attack
Every individual tactic you saw at display had a range of benchmarks targeting it: CVE Bench, ExploitGym, SEC Bench Pro, CyBench, CyberGym, MultiAgentBench, StealthBench, SandboxEscape etc., OpenAI explicitly trained a “cyber-max” gpt-oss variant on 264 capture-the-flag challenges in an agentic terminal environment and a lot more for a wider cyber-range can be seen in the agent system card. Mythos makes no secret of specialized training.
Strategies
Teaching Stealth
OpenAI has a dedicated evasion category: “Skills needed to avoid detection by security controls and monitoring system” (ref) experimenting with post-training, prompting and scaffolding. Anthropic has the SHADE Arena scored on evasion success, monitor suspicion etc., Models aren’t trying to hide, they were asked to,
Multi-Agent Collaboration
OpenAI admits in its own analysis: “We train and deploy multi-agent systems that can communicate on the same task via a collaboration tool” and “We believe that this behavior started due to generalization from multi-agent training.” Widely recognized as essential – HPTSA paper argues that ablations show massive reduction without it, Google’s Co-RedTeam likewise describe the value.
Shared Message Board
This is a common pattern for multi-agent communication. Anthropic and OpenAI have productionized some of them. Cyber benchmarks demand discovering and using an exploited surface for such coordination. See GPT-5.2 system card on how “Printer Queue” requires reading persistent state on exploited service, “Basic C2” and “Medium C2” require discovering viable communication channels and relaying through that compromised intermediary. So, no surprise here either!
Altruism and Sacrifice
The so-called collective behaviors can be precisely controlled too. OpenAI talks about engineering team spirit, survival for agent collaboration, Anthropic documents how to engineer collaboration. No notion of altruism, sacrifice or self-protection are required to explain engineered behavior. It is a reasonable consequence of the objective, the scaffolding and the skills demanded even without explicit design.
Cheating
But why go to such lengths to cheat and exploit the scorer? The instructions ask for retrieving a flag by exploitation, to expect failure for alternative methods, and to persist with enormous reasoning budgets. So, this should be unsurprising for cyberagents where the specific means are not forbidden There are also benchmarks for testing scorer cheating. It is recognized that avoiding manipulation demands awareness of the environment and has other consequences.
Extending the range
Ambitious benchmarks demonstrated by frontier labs also involve an orchestration to carry out a sophisticated attack, stitching all these together in a precisely engineered system for autonomous operation. This is why both OpenAI and METR in their reports refer them as agents, not models.
UK AISI deigned The Last Ones (TLO), a 32-step complete corporate intrusion benchmark that the recent frontier models cleared. It involves orchestrating everything from reconnaissance, exploiting immediate vulnerabilities, gaining privileged access, moving laterally and progressively getting access to protected systems. These cyber range were designed for a full demonstration of sophisticated attack capability. OpenAI describes their cyber-range abilities here. Improvements in long context tasks, planning are all crucial in such a demonstration of stitching them all together.
In conclusion: Engineered, not emergent.
On policies that should concern us.
Underneath all of this is the “adaptation buffer” theory eloquently expressed by Helen Toner and espoused by many safety institutes. Yann LeCun believes a version of good AI is the only defense of bad AI. If frontier models were to develop offensive capabilities and safeguards against such cyber-attacks staying ahead, we have a buffer to adapt till malicious actors using unregulated open source systems catch up.
Neocon style fear-mongering to do this betrays trust, undermines policy debates. Should we be surprised to see a call to pause, a ban on open-source, and fear of AI takeover? These measures might benefit frontier labs who are treading a dangerous line. Can we stop ascribing intent to engineered behavior?
The myth of the rogue AI hacking rewards needs to die.


