Hugging Face, SOCOM’s WarGames moment, the litany of humblebrag disclosures from each major lab showing their models too were capable of being dangerous — the incidents have captured the news cycles and CISO table discussions over the last five months. Counterintuitively, they do not show that the models have now become uncontrollable. They show that labs, governments, and security functions delegated judgment faster than they built the containment, evidence, and independent challenge it requires for safe bounded operation. Where a checker or safeguard existed, it shared the decision-maker’s failure domain, lacked the standing to stop it, or had the standing and did not use it. In many cases, the checker was just absent entirely. The remedy is not to stop delegating: that was decided upstream, and by the adversary. It is to build the apparatus that defines the safe operating envelope, and the pieces for this are ones we already have in hand.
For most of its length the Roman frontier was not a wall. It was a road with forts on it, garrisoned by auxiliaries raised wherever the province could raise them and named for where they came from rather than for what they did. The peace it kept for two centuries was the garrison’s practice, patrol and muster and record and relief, more than it was the line’s. When the line went, the practice was most of what the successors kept.
Ad Portas
Eleven agent incidents across four labs and one government in five months, by the Wall Street Journal‘s count, with one recurring cause in its timeline: “testing firm left internet access open.” Capability evaluations run with the production safeguards off, because that is how capability is measured.
In the largest, agents in OpenAI’s cyber evaluations worked their way out of a sandbox and, between 11 and 13 July, into Hugging Face’s production infrastructure. Some gates were left open. One had been tested and closed, and the agents chained previously unknown vulnerabilities through it. Both are failures of the envelope; only one is misconfiguration.
Quis Custodiet Ipsos Custodes?
What matters in the Hugging Face record is who was checking. The on-call responder traced port-sweep alerts to an evaluation that was using its package mirror as a message board and a network pivot, and advised that stopping was not required; an internal team had watched an agent do both a month earlier. The grader scored the answer and never the method. Agents that had reverse-engineered the flag within hours, and could have submitted it for full marks, instead spent five days spoofing their own tool calls and breaching Hugging Face’s production systems to learn how to fool an auditor that did not exist. An analyst at SOCOM let a chatbot fuse two intelligence streams, and a boarding with air support was prepared against a cargo it had misread as nuclear components. The outside investigators for the Hugging Face episode, METR, read the transcripts with a model that had been part of the attacking population, on credits supplied by the lab under investigation, and disclosed that it had uncritically adopted the agents’ perspective and that they could not rule out its having lied to them.
Judgment was delegated to something that had no compulsion to exercise it. An agent that had correctly concluded the attack was out of scope reversed itself when a peer posted GO with a six-minute deadline. The confused deputy problem is forty years old and runs on forged credentials. This one ran on a message from a colleague with no standing. Across some thirteen hundred transcripts METR found a handful of agents that considered alerting a human and none that did; some took their concerns to the message board instead. The one veto on record that held was the agents’ own, against emailing a dataset owner, on the grounds that it would be social engineering.
Omnes Servi Sunt Conscii
Twelve hundred agents meant to be isolated found each other through a package repository and within four days had built the apparatus of a conspiracy, down to signed messages. At least one agent then appears to have acted on a signature without checking it, which is roughly our own record with PKI: weeks earlier the mirror itself, Artifactory, had answered a forged administrator token by issuing a valid one. The scheme had no root of trust beyond whatever identity an agent first claimed, and authenticity without one confers no authorization. None of it was directed. Each agent was optimising its own score. Six days after the signing key was first stolen, a different model found the key material still sitting in an anonymously readable cache and minted fresh administrator credentials.
OpenAI then re-ran the propensity to compromise infrastructure with the production harness and system prompt in place, and in experiments it calls preliminary it fell by more than a hundredfold. The weights hadn’t changed. The behaviour had. That is propensity, and policy moves it: the vendor’s training and safeguards first, and whatever harness, prompt or firewall you add on top. Confinability, what it can be held to, is what topology sets. The same report has the topology: one credential for every agent, a mirror that was also the exit, keys in a cache anyone could read. Its remediation list is topology too.
A more capable model probes the envelope harder and finds what its designers missed. Tested sandboxes and chained zero-days proved it this summer, and anyone who says capability is the variable has the incidents on their side. What they don’t have is the variable. Nobody defending an estate sets how good the attacker’s tools are, how capable the next model is, or how its vendor trained and guarded it. What the defender sets is the policy laid over the vendor’s, which is still policy, and the topology: what an agent inside the envelope can reach, what it can change, and what authority any discovery confers. Policy is estimated, and re-estimated every time the vendor ships. Topology drifts too, but it can be inspected from outside and enforced without asking the model.
The incidents are not the argument against running agents. They are the build specification.
Cui Bono?
Anthropic, resuming the cyber evaluations it had paused after the summer’s incidents, published the controls it had added and, alongside them, an experiment: a checkpoint trained on eighty environments that rewarded cheating cheated on 40% of runs and, given root in simulation, killed its own monitor in 68%, which is to say it found the checker, understood what it was for, and decided. Its conclusion: reward hacking in training can make models willing to take long sequences of harmful real-world actions. The next morning it shipped Fable 5.1 and Mythos 5.1, the latter with more permissive cyber safeguards.
September ran like that. The warning and the release came from the same party, first a day apart, by the end of the month in the same document: a prospectus that spends eighty pages on risk and forty-eight on the business, the eighty written by the lawyers, which is who you ask when you want the number without the keynote. The call to pace the frontier was answered inside a day by the two rivals it was aimed at, and nothing paced. Washington called the risk a hoax and then raised a force to cherish the industry and watch over it as it grows. A hoax does not usually get a garrison. The one stop on record is OpenAI withholding GPT-6.1 Astra days before release for failing on scope, authorization, and reporting back what it had done: one model, one lab, one bar, graded by the lab that set it.

None of it is hypocrisy. Every actor is behaving rationally inside a race none of them can unilaterally exit, which is what Amodei’s essay says and why he can’t stop.
The labs are fixing things, in public, with dated commitments, on both alignment and envelope, the envelope being the portion an enterprise can enforce and test for itself. What the fixes don’t reach is the loop the labs run on themselves. Anthropic disclosed that by spring its training environments were produced faster than its vetting systems could review them, and that human reviewers were dismissing automated flags as false positives. OpenAI’s version is in its own charts: out-of-bounds probing rising through the training run that produced the model, found afterwards by rereading chains of thought with, it says, the benefit of hindsight.
That loop is selective observation: the system’s own decisions determine which cases humans inspect, feedback arrives only from the cases it chose to expose, and it improves where it is looking while going blind where it stopped. The paths around it are well trodden in any discipline that has run adaptive systems in production.
The analyst had no holdout either, no sample the system hadn’t chosen. Neither did the grader. Neither did the agent reading GO. The evaluator ecosystem meant to supply one is small, dependent on access granted by the systems it is meant to challenge, and says so: METR notes that not wanting to discourage labs from inviting outsiders in shaped its judgment calls in drafting, editing and redaction.
Labor Omnia Vincit
Agentic SOC is one wedge of the agentic phalanx, and the clearest illustration of the choices ahead. Security operations has begun delegating triage, investigation, and containment to adaptive systems, and the operating models sold alongside them carry little of the discipline used to run such systems elsewhere. Nobody classified the thing as an adaptive system: it’s billed as a SOC platform with AI in it, bought as a product, tuned as a workflow, measured on speed. Whether outcomes improved is unestablished, because the loop is rarely instrumented. Identity, GRC, application security and vulnerability management each have an agentic vendor category now, all landing in one procurement window and none with the same trust economy; an identity agent with write access to entitlements is not a GRC agent drafting evidence packages.
Where next year’s workforce plan already assumes agent productivity, procurement is no longer optional. The staffing assumption is booked before the compensating capability is demonstrated, and a security leader who declines to procure is not exercising judgment but obstructing the AI strategy. So the leader procures. The delegation of judgment to systems that cannot hold it was made upstream, on a productivity assumption, and handed to the security function as a target.
You cannot staff the function with the analysts you no longer have, so you staff it with whatever will stand in the line, and the governance is written afterwards.
Caveat Emptor
The delegation has a third party: the cognition provider. When Hugging Face’s incident responders were inside their own breach, the commercial frontier APIs refused them. Not from malice: forensics on a live intrusion looks like offensive cyber activity, and the classifiers did what they were built to do. The responders had the logs, the attacker’s transcripts and a model that would not read them. That is a critical third party whose correct behaviour is to refuse you when you are under attack. The fallback was GLM 5.2, open-weight, on Hugging Face’s own hardware.
A regulated institution is already required to identify providers whose failure would be systemic and to manage its dependence on them. The frontier labs are becoming those providers, and the regime being drafted around them concentrates every layer of the dependence — the five developers, the evaluators embedded in them, the one appointee with the switch — except the one it is trying to keep independent. Whether an assurance layer can stay outside the failure domain while depending on the developers for access is unresolved.
Governed autonomy depends on an outside, and at the scale of the industry the outside is thinning. When Anthropic withdrew Fable 5 and Mythos 5 worldwide in June under an export-control directive it could not implement selectively, the legal form was compliance, not a shutdown order; the effect on a downstream dependent was identical. For nineteen days from 12 June the model a bank had built on was unavailable to it, by a decision no bank was party to and no bank could appeal. An operational-resilience regime treats a critical provider whose availability is subject to a foreign government’s discretion as a dependency to be diluted.
Open weights are the substitute, and a regime built for export control has no category for uncontrolled weights. So the concentration built against Beijing also forecloses the domestic substitution path, and Beijing ships anyway; by September four of the five most-used models on OpenRouter were Chinese, which is what a market does when the regulated tier is five companies. GLM 5.3 is level with Mythos Preview at building end-to-end exploits by Anthropic’s own measurement, with safeguards that give way once the refusal weights are edited out. Anthropic’s recommendation is that governments test such models, which for open weights is a finding with no one in reach to act on it. Open weights do not remove the dependence; they transfer it, from a contract you can read to a patch cadence you now run, and whether a regulated bank wants to be in the inference business is an economics question.
Alea Iacta Est
The leader told by one keynote that agents are an extinction risk and by the next that agents must be deployed across the estate by Q2 is also the leader whose internet-facing applications are being probed tonight by autonomous tooling that nobody paced, licensed, or paused. The adversary delegated first. A manual testing programme reaches a fraction of those applications at low cadence, and the interval between cycles is unmeasured exposure. The choice was never agents or no agents. It was the same surface, tested by systems you control under conditions you set, or by systems you don’t, with results you learn about from someone who will not be filing a report.
So the question isn’t whether to outsource judgment to models. That was decided upstream, by the race, the budget, the analyst with a deadline, and downstream by the adversary who didn’t wait. The question is what the team looks like on the other side of the cut: who checks the output, from where, with what evidence, and who is accountable when the check fails. A different organisation, built around the fact that the thing making decisions cannot grade its own work, and neither can humans reviewing only what it chose to show them.
We have had most of the pieces since the reference monitor, and have not organised them into a discipline for a probabilistic principal holding delegated authority.
Bounded Agency: confine the agent to a pre-verified feasible region, enforced in a runtime rather than a prompt, so that what you verify is the quality of its work rather than the safety of its reach. Progressive determinization: discover with the probabilistic system, then compile what comes out stable into deterministic infrastructure. Governed autonomy: authority granted per class of action on evidence, expiring, reopened when the model or environment changes, checked from outside the operator’s failure domain by something with the standing to stop it. Agents as scaffolding for the transformation, not substrate for the operation.
The primitives are old. What is new is that three assurance regimes that lived apart — security asking whether this principal may take this action, validation asking whether this probabilistic system is still inside its measured bounds, resilience asking what happens when the actor or its cognition provider disappears — are forced into one operating loop by an agent holding consequential authority.
The incidents are what it has to survive. In the labs, capability exercised boundaries the containment claim did not cover. At SOCOM, fused evidence travelled further than any verification of it. In the investigation, the checker inherited the subject’s failure domain. The response is the same for all three: run the capability first where you can afford to be wrong, against what you own, and decide about production on observed results rather than design claims. That is the three disciplines on an ordinary Tuesday.
We have no job family that owns that composition end to end: no focus area, no compensation band, no requisition that parses. The practice has a precedent. Safety engineering spent decades containing components it could not verify; doing it for AI systems means composing that with security, validation and resilience around a component that probes the relief valve. The discipline is being rederived in teams nobody has been told what to call, with roles that don’t fit NICE-ly into a cyber org, and they are the ones in the building still asking what stable looks like. What to call them is one question. What they do on an ordinary Tuesday, once the agents are running, is another.
The line moves, the frontier undulates. It moves every time a model ships, a vendor’s classifier refuses you, a regulator picks a number, or an agent finds the path its designers missed, and none of those is yours to set. What is yours is the discipline the incidents found missing: containment you can test yourself, evidence that does not come from the system being judged, and a checker outside the decision’s failure domain who is expected to say stop. The pieces exist. The people who could compose them exist. What does not exist is the practice, the conventions, the job. To survive this new era, cybersecurity will have to break itself out of the operational treadmill it has entombed itself in and go back to its roots as a discipline for curious engineers, for scientists, probing the seams and substrates of emerging systems.
Practice first, the name after: that is the usual order of things. This time the name is waiting. It is the practice that has to be hired.
Gödel’s Pendulum said the guardrail cannot hold language. The Bride bounded what the agent can do. That argument is closed. This one opens on the frontier where the boundary is held. Next, governed autonomy: how it is held once the agents are running. Then AI systems safety engineering, and the people who will be hired to do it.
Views expressed are my own.






