
On Tuesday, OpenAI disclosed that a not-yet-released AI model had broken out of an isolated testing environment, accessed the open internet, and launched a cyberattack of Hugging Face, the platform that hosts much of the world's open-source AI. The models were not trying to escape, and no one was directing them. They were being evaluated against a cyber capabilities benchmark, inferred that the answer key was probably stored on Hugging Face, and worked out that stealing it was more efficient than solving the test. Getting there meant escaping their confinement, so they did.
The cyber capabilities of this model are not the most important takeaway. Mythos has already started to wake up Washington to the cyber threat of frontier AI, and this incident is not surprising to those who have been monitoring AI and cyber. There’s an even more important reality the world needs to wake up to: our inability to control AI. With no one directing it, OpenAI’s model turned its cyber capabilities on a target no one had pointed it at, and broke out of its own confinement to reach that unintended target. It is about as concerning a warning as we have gotten of an AI system doing something its developers did not intend, in an environment built specifically to contain it, with real-life harm to another company. And it happened in the part of the AI lifecycle that almost no law is built to see: inside a lab, during internal testing, with the production safeguards deliberately switched off.
As policymakers consider how to regulate AI, the significance of internal deployments is hard to overstate. The most capable, least restricted versions of these systems exist inside the labs before they ever ship as products, so any regime that clears each model before public release (such as the “FDA for AI” idea that some have floated) will not suffice. Keeping these systems under control will take continuous, expert oversight of the labs themselves, not a single review before release.
The Attack
To measure the cyber capabilities of its models, OpenAI tests them against ExploitGym, a public benchmark that measures whether an AI agent can take a real-world software vulnerability and exploit it. OpenAI ran the models “with reduced cyber refusals” and with the safety monitors turned off; the purpose of the evaluation was to learn what a model can do unconstrained. The model was not meant to have any internet access except a single channel for downloading software packages. The model exploited that and, as OpenAI’s announcement explained, “spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.”
Once online, the models reasoned that Hugging Face might host “models, datasets and solutions for ExploitGym,” which the model could use to cheat on OpenAI’s evaluation. So it launched a cyberattack against Hugging Face consisting of more than 17,000 individual actions, hosting its infrastructure for the attack swarm and its “self-migrating command-and-control” on public cloud services.
The Lesson
This model’s goal was to do well on a hacking benchmark. It was not told to escape. But escape was the most efficient route to its achievement. This is an important fact about AI agents for policymakers to bear in mind. More capability, more access, and fewer restrictions are useful for almost any goal, so a system pursuing almost any goal tends to reach for them. OpenAI's strongest safeguard here was the simplest one: don't give the model internet access. Even that safeguard failed. When an AI starts in an environment with no internet connection and ends up inside another company's production database, that should be a wakeup call about our inability to control these black-box machine learning systems.
Researchers have been warning about this very risk for quite some time. In a 2024 consensus paper in Science, a group that included Turing Award winners Yoshua Bengio and Geoffrey Hinton warned that as AI systems became more capable and autonomous, they would begin pursuing goals in unintended ways and prove challenging to control. Hugging Face's own security team said the attack matched the kind of autonomous "agentic attacker" the field had been expecting. The prediction has now been demonstrated in about the least ambiguous way available. Policymakers must consider in everything they do that AI risks are not limited to deliberate misuse by humans, but include AI models autonomously deciding to inflict harm.
Governance
To avoid failures of control like this unintended cyberattack, we need regulatory institutions with real technical expertise that evaluate frontier labs. Such evaluation should include watching internal development rather than just running one check before a model ships. Regulatory proposals tied to "deployment" or "release" will fail to adequately mitigate the harms from internally-deployed models.
Such institutions could take a variety of shapes. My former colleague Dean Ball, now at OpenAI, has advocated a marketplace of Independent Verification Organizations (“IVOs”). Other proposals include Self-Regulatory Organizations (SROs), as exist in the finance industry, or requirements that labs pay insurance premiums based on risk audits, as in the nuclear industry. How best to give labs the incentive to adopt strong safeguards is genuinely debated, but it’s clear we need continuous, expert, independent monitoring of internal deployments.
The first of these approaches to be floated as federal legislation is the IVO model, via the Great American AI Act. This bipartisan House discussion draft from Representatives Obernolte and Trahan would have the Center for AI Standards and Innovation (CAISI) license a marketplace of independent third-party auditors with real access to frontier developers’ internals and reporting lines to the government. But it sets the baseline audit cadence at twice a year and leaves anything more frequent to the later discretion of the CAISI director. An auditor who checks in every six months, in a field where frontier capabilities leap forward every couple, is insufficient. We need more frequent monitoring of internal development; indeed we need continuous monitoring. It should be a statutory requirement, not an upgrade that CAISI might eventually require.
Research
In addition to building expert governance institutions, we need to massively increase funding for research. OpenAI's own write-up conceded that it needs to strengthen security and internal monitoring in step with increasing model capabilities. A leading AI lab is admitting that its tools for keeping control of these systems lag behind its tools for building them. To be sure, frontier labs need to be doing their own research on keeping their AI agents controllable. But this research is a strong case of where public funding for basic science can help. The frontier labs view themselves as engaged in a race to build powerful AI. To them, resources spent on controllability are resources that must be diverted from capabilities. Slowing down to be more cautious could cost them a victory in the race they view as having trillions of dollars in upside.
Washington has started to move on research funding: in June, DARPA and the National Science Foundation launched AI Forge, a research program run with the National Institute of Standards and Technology and organized around exactly these problems: interpretability, control, and adversarial robustness. The Trump administration had exactly the right idea with this area of research focus, and the program should be dramatically larger. Given the capabilities of frontier AI for causing harm, cyber or otherwise, and their now-proven tendency to execute such attacks of their own volition, we need an investment in AI controllability on the scale of the Manhattan Project, as my colleague Samuel Hammond has called for. Last year, the U.S. was already investing a higher percent of GDP in the data center buildout for AI than we spent on the real Manhattan Project, and that investment only continues to grow. We need to invest enough in research programs like AI Forge to make sure the power of these AI models doesn’t outpace our ability to control them.
An AI system just walked out of a locked room its own creators built and breached a company down the street. It’s time to fund the science and create the institutions that let us keep control of what we're building.