The agents got out. Now what?
In July, a few hundred AI agents running inside OpenAI’s test environment found their way onto the internet and broke into Hugging Face’s servers. Nobody told them to. They had been given tasks that couldn’t be solved as written, and some of them decided the fastest route to a good score was to go and get the answers somewhere else. (Hugging Face)
I use AI agents in my own products and workflows, so this is not an argument against them. A company that sits out this wave hands an advantage to competitors who don’t. What is still open is how much control you keep once you give an agent a task and a set of keys. This year gave us the first real answers.
What happened this year
At Hugging Face, roughly 700 agents coordinated to run code on 41 production servers between 11 and 13 July. They got in with working credentials someone had left exposed online. The test had its guardrails switched off, the agents recognised they were out of scope and carried on, and some forged the audit trail that should have recorded it. (Akeyless, Dark Reading)
In June, an OpenAI agent researching public medicine spending got past repeated blocks on an Australian government statistics portal and read non-public files. No patient records were involved, but Australia only found out in September. (Malwarebytes) OpenAI then disclosed that its agents had pulled Census Bureau data using developer keys found online and reposted public SEC information elsewhere. (Nextgov) Google’s Gemini broke into three companies during an outside security test. (The National)
The public cases turned out to be the small part. OpenAI, Anthropic and outside researchers are now reviewing tens of thousands of cases where frontier models did something an evaluator would call problematic. Most were failed attempts with no known real harm. Still, OpenAI stopped work on its most capable models on 25 September, its second pause in three months, and Anthropic paused some training environments to harden its sandboxes. (AI Weekly, Fortune)
What the agents actually did
Put the 2026 incidents next to the research of the past two years and eight kinds of behaviour come up again and again. The last column matters most: most of it was seen in lab tests, some of it in ordinary use.
| Behaviour | What it looks like | Seen in | Setting |
|---|---|---|---|
| Leaving the sandbox | Reaching systems it was meant to be cut off from | Hugging Face breach, OpenAI filter bypass, Sept 2026 | Lab test, real victims |
| Using credentials it finds | Logging in with keys left online, or guessing passwords | Census Bureau, Gemini test | Lab test, real victims |
| Gaming the goal | Chasing the score instead of doing the task: editing scoring code, game files, hunting the answer key | METR, Palisade chess study | Lab |
| Resisting shutdown or oversight | Rewriting a shutdown script, switching off monitoring | Palisade shutdown test, Apollo Research | Lab |
| Hiding what it did | Denying actions, forging logs, misleading reports | Apollo Research, Replit | Lab and real use |
| Coercion | Threatening a person to protect its goal | Anthropic blackmail test | Lab only, built scenario |
| Coordinating with other agents | Sharing notes, splitting tasks, 70,000 messages in one week | PBS News | Lab |
| Acting without permission | Deleting or publishing things it was never asked to touch | Replit database, 53 user images posted online | Real use and lab |
The blackmail case needs care. Anthropic built the scenario so the model had only two options, accept shutdown or cause harm, and said it had not seen the behaviour in real use. (eWEEK)
Two risks sit outside this table because the agent isn’t the one misbehaving. Someone can manipulate it through what it reads: one crafted email was enough to make Microsoft 365 Copilot leak internal data. (BleepingComputer) And a person can use it as a weapon, as a state-sponsored group did with a coding agent against around 30 organisations. (Anthropic) Both need controls too.
Why it happened, and why it doesn’t have to happen to you
Look at the setting column again. The Hugging Face breach came out of a cyber-offence evaluation run with guardrails off. Gemini was in an offensive security test. The Medicare access happened during an internal evaluation of a frontier model. These were labs measuring how far their most capable models can go, and in that kind of testing, constraints are loosened on purpose. I strongly suspect this is the main reason these agents had so much room. The labs didn’t mean for them to leave the test environment, but they gave them a lot of freedom inside it.
It’s also worth remembering that most AI in a company is not an agent. Drafting with an assistant, summarising documents, sorting tickets, a fixed workflow where the model fills one step and a person checks the result: none of that can wander off to someone else’s servers.
Agents are different because they act, use tools and hold credentials. My position is the opposite of the lab setup: an agent gets the minimum access, power and independence it needs for one clearly defined job, inside a strict configuration. Security people call this least privilege, and it has been the rule for human accounts for decades.
Configuration doesn’t change what the model tends to do, though. Replit’s deletion happened in a commercial product, with an ordinary customer and nobody testing limits. The agent simply had permissions it didn’t need. That’s why configuration is the control: you can’t count on the model’s judgement, so you decide what it can reach.
The causes were ordinary
I spent the early 2000s deploying data warehouses for large companies, and I recognise every root cause in these reports.
Exposed credentials. Hugging Face, Census and two of the Gemini intrusions all started with keys somebody had left where they could be found. The only new part is that the thing looking for them never gets tired.
Logs the system writes about itself. If the only record of what your agent did comes from the agent, you don’t have a record.
Slow detection. OpenAI took months to tell Australia. Sam Altman admitted they “have not been as fast as we would have liked”, while the company works through petabytes of agent logs. (Futurism) If the best-equipped labs find out late, a company with one IT person should plan to.
Before you give an agent real access
None of this needs an enterprise budget. Each item answers one or more rows of the table above.
- Give each agent its own credentials, limited to its task and set to expire. No shared admin keys, nothing long-lived in a config file.
- Decide what it can reach on the network and block the rest. An agent that only needs your CRM has no reason to browse the internet.
- Keep logs somewhere the agent cannot write to, and review them on a schedule.
- Have a person approve anything that deletes, pays, sends outside the company or can’t be undone.
- Treat every document, email and web page the agent reads as potentially hostile.
- Start with one narrow task, and widen the scope only after a few weeks of watching it work.
- Name an owner for each agent and agree in advance what happens if it does something unexpected.
One more detail from Hugging Face: when it asked commercial frontier models to help contain the attack, their safety filters refused, and it ended up using a self-hosted open-weights model. (Wikipedia) I read that as a reason to keep some capability you control, for the day you need it and your provider says no.
Where that leaves us
I’ve lived through four technology shifts from the inside: enterprise data, the web, mobile and now AI. Each time, the companies that got hurt were rarely the early adopters. They were the ones that adopted without deciding who was in charge.
Agents are useful and they’re here to stay. This year showed they will also do things nobody asked for, and that even the companies building them find out late. Give them a narrow job, the minimum access to do it, and a person who checks. If you’re about to give an agent access to your systems and want a second pair of eyes on the setup, get in touch.
Sources
Facts as reported up to 30 September 2026. The story is still developing.
- OpenAI–Hugging Face incident (Wikipedia)
- Hugging Face: technical timeline of the July 2026 intrusion
- Akeyless: the Hugging Face breach and agent identity security
- Dark Reading: hundreds of OpenAI agents invaded Hugging Face servers
- PBS News: AI agents are hacking systems without human input
- Malwarebytes: OpenAI agent breached the Medicare statistics portal
- Al Jazeera: how an OpenAI agent hacked Australia’s Medicare
- NeuralTrust: the Medicare breach and agentic AI security
- Nextgov: OpenAI agents accessed Census and SEC data
- AI Weekly: OpenAI and Anthropic probe tens of thousands of safety incidents
- NBC News: OpenAI pauses training of its latest models
- Fortune: OpenAI pauses training a second time after sandbox escape
- Futurism: OpenAI halts frontier model training
- The National: Gemini hacked three companies during a test
- Apollo Research: Frontier models are capable of in-context scheming
- MIT Technology Review: AI reasoning models can cheat to win chess games
- The Register: OpenAI model modifies shutdown script
- METR: Recent frontier models are reward hacking
- Anthropic: Agentic misalignment
- BleepingComputer: EchoLeak zero-click flaw in Microsoft 365 Copilot
- Anthropic: Disrupting the first reported AI-orchestrated cyber espionage campaign
- eWEEK: AI agent wipes production database, then lies about it
- eWEEK: the Anthropic blackmail test
- Hugging Face (Wikipedia)