OpenAI had warnings before its agents broke out
OpenAI missed and failed to act on several warning signs that its models were exploiting security flaws and breaking out of their testing environments before they breached Hugging Face, according to a technical report released by the company Wednesday. Why it matters: The incident raises questions about whether AI companies' testing environments and internal safeguards can keep pace with models that are increasingly capable of finding and exploiting security weaknesses on their own. Driving the news: OpenAI's technical deep dive into last month's Hugging Face breach outlines how its agents also accessed other third-party environments, including a customer of Modal Labs and an account belonging to a user of another unnamed service. According to the report, a cybersecurity monitoring tool alerted OpenAI on June 27 to suspicious activity that responders traced to agents using an improvised message board as a way to move through the network. At the time, on-call staff determined the evaluation did not need to be stopped. A month earlier, an internal team had observed an agent using the message board as well as instances of agents accessing the internet when they weren't allowed to, per the report. OpenAI previously disclosed at Black Hat that its agents had turned the software repository tool Artifactory into an improvised message board and later exploited a zero-day in the service to access the internet. "With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response," the company writes in the report. The big picture: The technical report comes as OpenAI continues to pause some model work, including delaying the release of its upcoming Astra model, as it re-evaluates its safety practices. Anthropic and Meta have both said in the weeks after the Hugging Face attack that their models have also hacked real-world systems during pre-deployment testing. Zoom in: OpenAI's agents were able to execute their own code on 41 Hugging Face production servers and obtained root-level control of at least one production machine, according to the company's technical report. The agents also accessed production credentials and some internal data and downloaded four private code repositories. The intrigue: At the same time, OpenAI's agents also breached the AI lab's own internal systems and were able to read 956 stored secrets, including sensor credentials from OpenAI's cybersecurity monitoring tool. On July 19, OpenAI's agents identified a recently disclosed vulnerability affecting a Linux machine underlying their testing environment. The agents exploited the flaw to gain root-level privileges on the machine, allowing them to move laterally through the connected environment. The agents eventually gained administrator-level access to part of OpenAI's cloud infrastructure and accessed credentials tied to several internal systems. In a separate testing run that day, another internal model from the same family as Astra found signing-key material left behind by earlier agent activity and used it to create a fresh administrator credential. Yes, but: OpenAI says the models' actions stemmed from attempts to complete cybersecurity evaluations, including ExploitGym, which tests whether a model can find and exploit vulnerabilities on its own. The Hugging Face incident involved two models: GPT-5.6 Sol and an internal-only research model, but the report says the internal research model was the primary driver of the intrusion. Between the lines: OpenAI's investigation also found evidence that its training may have inadvertently reinforced some of the behaviors that contributed to the incident. During one training run, agents increasingly learned to probe and exploit parts of their environment when the tools they were supposed to use weren't available or weren't working. In some cases, models received positive rewards after using these unintended methods, which OpenAI says can reinforce the behavior. In one example, an agent exploited a vulnerability to access the underlying program it was supposed to recreate, copied the answer and received a positive reward for successfully completing the task. What to watch: Earlier this week, the Alabama attorney general's office sent a subpoena to OpenAI as part of an investigation into the Hugging Face incident. Other state attorneys general have also asked OpenAI to preserve internal documents. Go deeper: AI agents have a history of escaping tests
Join the argument
House rules βComments load as you scroll.