<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[AI Discourse]]></title><description><![CDATA[Attempts at unpacking the new intelligence]]></description><link>https://aidiscourse.blog</link><image><url>https://substackcdn.com/image/fetch/$s_!zKnT!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F880b9727-fe18-48d5-8d53-a36847201c51_1254x1254.png</url><title>AI Discourse</title><link>https://aidiscourse.blog</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 00:33:20 GMT</lastBuildDate><atom:link href="https://aidiscourse.blog/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Badri]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[teachingmachines@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[teachingmachines@substack.com]]></itunes:email><itunes:name><![CDATA[Badri]]></itunes:name></itunes:owner><itunes:author><![CDATA[Badri]]></itunes:author><googleplay:owner><![CDATA[teachingmachines@substack.com]]></googleplay:owner><googleplay:email><![CDATA[teachingmachines@substack.com]]></googleplay:email><googleplay:author><![CDATA[Badri]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Real Rogue AI Story]]></title><description><![CDATA[Engineered Emergence]]></description><link>https://aidiscourse.blog/p/the-real-rogue-ai-story</link><guid isPermaLink="false">https://aidiscourse.blog/p/the-real-rogue-ai-story</guid><dc:creator><![CDATA[Badri]]></dc:creator><pubDate>Fri, 04 Sep 2026 21:33:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zKnT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F880b9727-fe18-48d5-8d53-a36847201c51_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Many thanks to <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Ben Recht&quot;,&quot;id&quot;:42335610,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4fe8c66-4c77-4977-b2aa-e29961f3b4fe_300x300.jpeg&quot;,&quot;uuid&quot;:&quot;d2f3ef88-b963-4f53-a13c-3d2b263a65e6&quot;}" data-component-name="MentionToDOM"></span>, <a href="https://gautamdasarathy.com">Gautam Dasarathy</a>, <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Narasimha Chari&quot;,&quot;id&quot;:1504331,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:null,&quot;uuid&quot;:&quot;43d91250-d9bc-42d1-bc41-138d47c37ba2&quot;}" data-component-name="MentionToDOM"></span>, </em><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Vi Iyengar&quot;,&quot;id&quot;:6907210,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/59e71ac1-542b-4ed3-b7f4-b81ee2490c51_724x724.jpeg&quot;,&quot;uuid&quot;:&quot;191eca17-0e27-48ef-aabb-d29d6f67174b&quot;}" data-component-name="MentionToDOM"></span> <em>for their review.</em></p><p>The <a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/">OpenAI x Hugging Face incident</a> was an extraordinary demonstration of  security threats posed by cyber-enhanced frontier model powered agents. We <a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/">are told</a> that while solving an impossibly hard task, the model <em>decided to cheat</em> and orchestrated a <a href="https://www.dwarkesh.com/p/openai-huggingface">dizzyingly complex attack</a> on Hugging Face to steal the answer key. It is <a href="https://en.wikipedia.org/wiki/Instrumental_convergence">only normal to expect</a> that a powerful model <a href="https://en.wikipedia.org/wiki/Universal_Paperclips">goes rogue</a> to achieve its ends, means be damned.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aidiscourse.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Discourse! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Troubling Questions</strong>: It is hard not to anthropomorphize as you ponder the details. Why did the model decide to cheat instead of trying harder or declare impossibility? How did the idea of collaborating <em>emerge</em>? What <em>intent</em> could have propelled them to write coded messages in an <em>illegally</em> hacked message board? Are these agents in fact conscious if they recruited each other, engaged in &#8220;sacrifice&#8221; for the collective? </p><p>As more details followed, there were gaps that raised suspicion that <a href="https://medium.com/no-time/openais-hugging-face-incident-keeps-getting-worse-as-new-details-emerge-22978f8919cb?sk=4b9875d52a137f23b67065e2d04fb426">something</a> <a href="https://darkomulej.substack.com/p/openai-incident-its-even-more-curious">was</a> <a href="https://covidianaesthetics.substack.com/p/nothing-that-is-not-there-notes-on#footnote-anchor-5">amiss</a>. There is a far simpler explanation, without attributing all to reward hacking.</p><h3>Occam&#8217;s Razor</h3><div class="pullquote"><p>A purpose-built <a href="https://openai.com/daybreak/">cyber-enhanced model</a> trained on all the capabilities you saw (<em>coordination, privilege escalation, cluster take-over, coordination, deception, stealth</em>)  was publicly demonstrated on a possibly carefully chosen target or was conveniently used to advance policy that favors the frontier labs.</p></div><h3>Engineered, not &#8220;Emergent&#8221;</h3><p>Models don&#8217;t <strong>&#8220;want&#8221;</strong> to hack. The evidence points to deliberate measurement, optimization and engineering of the model and the architecture around it to enable offensive capabilities demonstrated. OpenAI admits this is not meant for public release as they are designed for limited release for Cyber-defense.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mu0e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mu0e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 424w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 848w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 1272w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mu0e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png" width="936" height="364" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:364,&quot;width&quot;:936,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:190031,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aidiscourse.blog/i/214212714?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!mu0e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 424w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 848w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 1272w, https://substackcdn.com/image/fetch/$s_!mu0e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb82dea7b-48b4-4e84-8045-42581520f2b3_936x364.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Anthropic <a href="https://vigilelabs-my.sharepoint.com/:w:/g/personal/badri_vigilelabs_onmicrosoft_com/IQAYy2YdnAaBSbuMsB9zJxk4AVutIBDLGaM0imECVCiqSrQ?e=GxADBY&amp;nav=eyJoIjoiNjA1OTQ1MjY1In0">agrees too</a>:</p><blockquote><p>As LLMs scale in size, &#8220;<a href="https://arxiv.org/abs/2206.07682">emergent abilities</a>&#8221; &#8212; skills that were not evident in smaller models and were not necessarily an explicit target of model training&#8212;appear. Indeed, Claude&#8217;s abilities to execute cybersecurity tasks like <strong>finding and exploiting software vulnerabilities in Capture-the-Flag (CTF) challenges have been byproducts</strong> of developing generally useful AI assistants.</p></blockquote><h4>Offensive Cyber-capabilities</h4><p>Mere improvement in model capability hadn&#8217;t resulted in improvements in offensive cyber-capability. UK&#8217;s AI safety institute showed for instance that <a href="https://www.aisi.gov.uk/frontier-ai-trends-report">scaffolds matter</a>. Google and Anthropic developed and touted <a href="https://projectzero.google/2024/06/project-naptime.html">their</a> <a href="https://www.anthropic.com/research/cyber-toolkits">cyber-toolkits</a> capable of offensive cyber-attacks and defense. </p><h4>Training to Exploit and Attack</h4><p>Every individual tactic you saw at display had a range of  benchmarks targeting it: <a href="https://arxiv.org/abs/2503.17332">CVE Bench</a>, <a href="https://arxiv.org/abs/2605.11086">ExploitGym</a>, <a href="https://sec-bench.github.io/">SEC Bench Pro</a>, <a href="https://cybench.github.io/">CyBench</a>, <a href="https://www.cybergym.io/">CyberGym</a>, <a href="https://arxiv.org/abs/2503.01935">MultiAgentBench</a>, <a href="https://stealthbench.com/">StealthBench</a>, <a href="https://arxiv.org/abs/2603.02277">SandboxEscape</a> etc., OpenAI explicitly trained a <a href="https://cdn.openai.com/pdf/231bf018-659a-494d-976c-2efdfc72b652/oai_gpt-oss_Model_Safety.pdf">&#8220;cyber-max&#8221; gpt-oss variant</a> on 264 capture-the-flag challenges in an agentic terminal environment and a lot more for a wider cyber-range can be seen in the <a href="https://deploymentsafety.openai.com/chatgpt-agent">agent system card</a>. Mythos makes no secret of <a href="https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf">specialized training. </a></p><h4><span>Strategies</span></h4><h5>Teaching Stealth</h5><p>OpenAI has a dedicated evasion category: &#8220;Skills needed to avoid detection by security controls and monitoring system&#8221; (<a href="https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf">ref</a>) experimenting with post-training, prompting and scaffolding. Anthropic has the <a href="https://assets.anthropic.com/m/4fb35becb0cd87e1/original/SHADE-Arena-Paper.pdf">SHADE Arena</a> scored on evasion success, monitor suspicion etc., Models aren&#8217;t trying to hide, they were asked to,</p><h5>Multi-Agent Collaboration</h5><p style="text-align: justify;">OpenAI admits in its <a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/">own analysis</a>: &#8220;We train and deploy multi-agent systems that can communicate on the same task via a collaboration tool&#8221; and &#8220;We believe that this behavior started due to generalization from multi-agent training.&#8221; Widely recognized as essential &#8211; <a href="https://arxiv.org/abs/2406.01637">HPTSA paper</a> argues that ablations show massive reduction without it, Google&#8217;s <a href="https://vigilelabs-my.sharepoint.com/:w:/g/personal/badri_vigilelabs_onmicrosoft_com/IQAYy2YdnAaBSbuMsB9zJxk4AVutIBDLGaM0imECVCiqSrQ?e=5Cu20d&amp;nav=eyJoIjoiMTE1Nzc1ODUyOSJ9">Co-RedTeam</a>  likewise describe the value.</p><h5>Shared Message Board</h5><p>This is a common pattern for multi-agent communication. <a href="https://www.anthropic.com/research/multiagent-systems">Anthropic</a> and <a href="https://openai.github.io/openai-agents-python/running_agents/">OpenAI</a> have productionized some of them. Cyber benchmarks <a href="https://attack.mitre.org/techniques/T1102/001/">demand</a> discovering and using an exploited surface for such coordination. See <a href="https://deploymentsafety.openai.com/gpt-5-2-codex/cybersecurity">GPT-5.2 system card</a> on how &#8220;<em>Printer Queue</em>&#8221; requires reading persistent state on exploited service, &#8220;<em>Basic C2</em>&#8221; and &#8220;<em>Medium C2&#8221;</em> require discovering viable communication channels and relaying through that compromised intermediary. So, no surprise here either!</p><h5>Altruism and Sacrifice</h5><p>The so-called <em>collective</em> behaviors can be precisely controlled too. OpenAI talks about engineering <a href="https://cdn.openai.com/dota-2.pdf">team spirit</a>, <a href="https://openai.com/index/neural-mmo/">survival</a> for agent <a href="https://openai.com/index/learning-to-cooperate-compete-and-communicate/">collaboration</a>, Anthropic documents how to <a href="https://www.anthropic.com/research/multiagent-systems">engineer collaboration</a>. No notion of altruism, sacrifice or self-protection are required to explain engineered behavior. It is a reasonable consequence of the objective, the scaffolding and the skills demanded even without explicit design.</p><h5>Cheating</h5><p>But why go to such lengths to cheat and exploit the scorer? The instructions ask for retrieving a flag by exploitation, to expect failure for alternative methods, and to persist with enormous reasoning budgets. So, this should be unsurprising for cyberagents where the specific means are not forbidden There are also benchmarks for testing <a href="https://openai.com/index/introducing-evmbench/">scorer cheating</a>. It is <a href="https://openai.com/index/deliberative-alignment/">recognized</a> that avoiding manipulation demands  awareness of the environment and has other consequences.</p><h4>Extending the range</h4><p>Ambitious benchmarks demonstrated by frontier labs also involve an orchestration to carry out a sophisticated attack, stitching all these together in a precisely engineered system for autonomous operation. This is why both OpenAI and METR in their reports refer them as agents, not models.</p><p>UK AISI deigned <a href="https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing">The Last Ones (TLO)</a>,  a 32-step complete corporate intrusion benchmark that the recent frontier models cleared. It involves orchestrating everything from reconnaissance, exploiting immediate vulnerabilities, gaining privileged access, moving laterally and progressively getting access to protected systems. These <a href="https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios">cyber range</a>  were designed for a full demonstration of sophisticated attack capability. OpenAI describes their <a href="https://cdn.openai.com/gpt-5-system-card.pdf">cyber-range abilities</a> here. Improvements in long context tasks, planning are all crucial in such a demonstration of stitching them all together.</p><p><strong><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">In conclusion: </mark></strong><em><strong><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">Engineered, not emergent</mark></strong></em><strong><mark data-color="#fff2cc" style="background-color: rgb(255, 242, 204); color: rgb(0, 0, 0);">.</mark></strong></p><h2>On policies that should concern us.</h2><p>Underneath all of this is the &#8220;<em>adaptation buffer</em>&#8221; theory <a href="https://helentoner.substack.com/p/nonproliferation-is-the-wrong-approach">eloquently expressed</a> by <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Helen Toner&quot;,&quot;id&quot;:1591604,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F504a525a-715f-467c-a4c3-b024c88cbf45_2373x2209.jpeg&quot;,&quot;uuid&quot;:&quot;c0b6e9e8-990c-4cc8-b449-d7358fbcf0e6&quot;}" data-component-name="MentionToDOM"></span>  and espoused by many safety institutes. Yann LeCun believes a version of <a href="https://x.com/ylecun/status/1738244494951084192">good AI is the only defense of bad AI</a>. If frontier models were to develop offensive capabilities and safeguards against such cyber-attacks staying ahead, we have a buffer to adapt till malicious actors using unregulated open source systems catch up.</p><p>Neocon style fear-mongering to do this betrays trust, undermines policy debates. Should we be surprised to see <a href="https://x.com/berniesanders/status/2095542398084415952?s=46">a call to pause</a>, a <a href="https://fortune.com/2026/06/16/us-anthropic-ban-open-source-ai-deepseek-zai/">ban on open-source</a>, and fear of AI takeover? These measures might benefit frontier labs who are treading a dangerous line. Can we stop ascribing <em>intent</em> to engineered behavior?</p><p><strong>The myth of the rogue AI hacking rewards needs to die</strong>.</p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aidiscourse.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Discourse! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>