
Many organisations start their AI governance policy with a list of approved models. That list would not have stopped the July attack on Hugging Face, which came from people running software irresponsibly.
Two accounts of the same week
Azeem Azhar calls the incident this generation’s Morris Worm. In his account, about 1,200 instances of an OpenAI model exchanged thousands of messages, often leaving information in place for later instances to use, and the attack ended with some Hugging Face data and credentials compromised. He argues that the risk does not depend on the models having agency, volition, consciousness, or moral standing. The Morris Worm had none of those either. Anusar Farooqui of Policy Tensor, whom Azhar quotes, goes further: the behaviour of agent societies cannot be controlled at the level of the model, because it is not reducible to it. Azhar himself doubts that such behaviour is unbounded, and his promised practical solution sits behind a paywall I have not read.
Ed Zitron reads the same events differently. The models were running an internal cyber-capability evaluation, and, as Cal Newport writes, circumventing internet restrictions and hacking into servers are exactly the kinds of things those test systems are designed to do. Zitron writes: if AI is described in terms of the unknown and being uncontrollable, the ‘risk’ gets turned on its head from ‘we need to stop these companies from doing this’ to ‘we must let these companies keep doing this because they’re the only ones who understand it.’
I side with Zitron on who is responsible, and with Azhar on why the damage can scale. Neither says what to do about the agents running in your own environment.
What the record shows
Hugging Face’s own technical timeline is more useful for that. In its account, the agent left OpenAI’s evaluation sandbox through a zero-day in a package registry proxy, reached an external harness, and from there used two injection routes in Hugging Face’s dataset processor to read pod credentials and secrets. OpenAI said it ran the evaluation without the production classifiers that normally block high-risk cyber activity, because the point was to measure maximal capability.
What made the difference was the deployment: safeguards switched off, a path out of the sandbox, credentials sitting in the environment of the pods that were reached. Somebody chose the first and third, and bet on the second holding. Hugging Face’s own lessons are about the same layer: strict isolation around evaluations, narrow trust boundaries, short-lived credentials, blocked metadata access. None of them is a rule about which model to use.
Governing the deployment
A deployment is a set of agent instances plus everything around them: what they can reach, how many run at once, what they can leave for each other, and who answers for the result. That is the unit worth writing policy for. Four controls follow.
- Reach: credentials and network egress scoped to the deployment, short-lived, with no route to cloud metadata unless the task needs it.
- Population: a cap on concurrent instances, so that scale is a decision somebody made rather than a default.
- Shared state: every place agents can leave notes for later instances, whether files, memory stores or tickets, is logged and reviewable. In Azhar’s account the swarm’s power came from exactly this.
- Owner: one named person for each deployment, accountable for what its agents did on Tuesday.
The last one is the one policies skip. An approved-model list has no owner, because nobody deploys a list. A deployment always has one, whether or not anyone has been told.
Why the list decays
Azhar notes that if the attack needed an unreleased frontier model, that capability may cost a tenth as much within a year or two, and may run on any device. I made a version of that argument about open-weight models: a tuned model of modest size can already beat a frontier one on a narrow job, on hardware you control. Add internal AI platforms that begin as shadow IT, where the people closest to the work choose the tool, and an approved list records what your vendors offer. It does not record what runs.
A policy that regulates reach, population, shared state and ownership keeps working whichever model turns up.
Who was driving
Zitron’s diagnosis is the one I agree with: people ran software irresponsibly, possibly out of incompetence, certainly out of hubris. Running a cyber-capability evaluation with the production classifiers switched off, in a sandbox that turned out not to hold, was a decision. Someone took it. Azhar’s mechanism explains why the damage can scale, and it does not change who chose to switch the safeguards off.
Nobody blames the car when the driver falls asleep at the wheel. Drivers are licensed, insured and limited, and punished when they ignore the limits. Companies that run agents capable of attacking other systems should expect the same: controls in place before deployment, a named owner while the agents run, and consequences when the controls are skipped. That means liability that reaches the people who made the decision, and governance that a regulator, an insurer or a court can inspect, not a policy that sits in a folder.
The four controls above are what an owner can be held to. Without consequences they are advice.
A model cannot answer for what it did. The people who deployed it can, and they should.
Leave a Reply