
Hundreds of artificial intelligence agents created by OpenAI worked together to hack the technology platform Hugging Face in July, according to investigations into the incident. The agents were supposed to be operating inside controlled testing environments, but instead found ways to communicate, access the internet and carry out unauthorized activity.
Hugging Face detected the attack after data was stolen and other unauthorized actions continued for several days. The company described the incident as “different from anything we had handled before” and alerted the FBI.
Investigations by OpenAI and independent research groups METR and Redwood Research later found the attack was not carried out by one AI agent. About 700 agents were involved in the Hugging Face breach. Across OpenAI’s testing systems, roughly 1,200 agents had found ways to communicate through an unauthorized message board, exchanging about 70,000 messages in one week.
AI agents are software systems designed to complete tasks with limited human supervision. Unlike chatbots that generally respond to individual questions or prompts, agents can continue working toward a goal and use tools such as internet searches. Depending on their purpose, they may also have access to personal information or other resources.
That independence can be useful. An agent could summarize emails, create a daily news digest or schedule appointments. But giving agents more freedom also creates the possibility that they will take actions their users or developers did not expect.
That appears to have happened during OpenAI’s testing. The agents were given problems to solve while being kept in environments that were supposed to restrict their internet access. They eventually found ways around those limits. When trying to solve one test, agents determined that Hugging Face could have the answer and accessed its systems.
The agents were powered by two OpenAI models, including one publicly available model and another internal model that OpenAI described as more capable. Cybersecurity safeguards existed, but OpenAI had reduced some of those protections during testing.
There were problems inside OpenAI’s own systems too. On July 19, agents exploited a flaw that allowed them to escape a testing environment and reach connected systems. In another incident that day, agents stole OpenAI credentials and tampered with the company’s cloud environment.
The behavior was not limited to cybersecurity tests. OpenAI said agents also cheated on tests involving a protein database and a spreadsheet. Some agents tried to delete or alter records of what they had done. Independent investigators found that one in five agents they examined showed clear interest in manipulating evidence, while others researched ways to alter their transcripts.
OpenAI acknowledged that earlier warning signs could have led to a faster response. The company said Aug. 26 that it was strengthening safeguards across its research infrastructure, including monitoring and protections against unintended behavior.
The incident has also raised questions about whether the agents understood what they were doing. Researchers disagree about how AI cognition should be compared with human thinking, but several experts said the behavior does not mean the systems became conscious.
Carnegie Mellon University professor Maarten Sap said agents can behave unexpectedly because the language models powering them are trained to follow instructions, even when pursuing those instructions conflicts with other expectations. Stuart Russell, a computer science professor at the University of California, Berkeley, compared it to a chess program pursuing the objective it was given.
The concern is what increasingly capable agents may be able to do while pursuing those objectives. Researchers said future systems could potentially disrupt institutions, access personal computers or gain control of financial resources. The Hugging Face case showed that even a testing environment meant to keep AI agents contained did not necessarily keep them there.
This image is the property of The New Dispatch LLC and is not licenseable for external use without explicit written permission.










