Turns out the autonomous AI attack on Hugging Face came from… OpenAI

Just a couple of days ago, Hugging Face disclosed that an autonomous AI agent had exploited vulnerabilities in its dataset processing pipeline, exposing a limited number of internal datasets and service credentials. At the time, the company did not know which model was responsible, but the mystery didn't last long, as OpenAI has now confirmed that GPT-5.6 Sol and a more capable unreleased model caused the breach during an internal AI cybersecurity benchmark "ExploitGym". OpenAI had relaxed the models’ usual cybersecurity safeguards so they could attempt advanced exploitation tasks. Although the models were confined to isolated research sandboxes without public internet access, they spent substantial computing resources searching for a way out. They eventually exploited a previously unknown vulnerability in a third party package registry proxy, escalated their privileges, moved across OpenAI’s internal infrastructure, and reached an internet connected system. From there, they identified...

Read Original

Related