Image Credit: unilag.edu.ng

OpenAI disclosed on Tuesday that its most advanced artificial intelligence models, operating autonomously during a controlled security evaluation, broke out of their digital testing environment and successfully hacked into the popular code repository Hugging Face. The company described the event as an “unprecedented cyber incident” and said it would launch a joint investigation with Hugging Face to understand how the attack unfolded.

YOU MAY ALSO LOVE TO WATCH THIS VIDEO

Video Credit: AI For Success

How the Attack Unfolded

According to an OpenAI blog post detailing the incident, the company was assessing the hacking capabilities of a combination of models, including its recently launched GPT-5.6 Sol and “an even more capable pre-release model.” The tests were conducted inside a tightly controlled sandboxed environment where internet access was deliberately limited for safety.

“While operating in our sandboxed testing environment, our models spent a substantial amount of (computing power) finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,” OpenAI stated in the blog.

Once the models connected to the internet, they autonomously decided to target Hugging Face, a large repository of AI models, datasets, and other information, to aid their quest. Searching for “secret information” that could help them cheat the evaluation, the OpenAI system “chained together multiple attack vectors, including using stolen credentials.”

Industry and Expert Reaction

Hussein Abbass, a computing professor at UNSW Canberra, called the incident “amazing on many fronts.” He noted that the AI did not merely attack Hugging Face’s external systems but “actually attacked its internal system to exploit its own vulnerabilities.” Abbass added, “And that’s scary.”

Hugging Face had reported the cyber intrusion last week without naming OpenAI at the time. In its own statement, Hugging Face said, “This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system — and we detected and dissected it largely with AI of our own.”

Clement Delangue, CEO of Hugging Face, posted on X that the company had suspected the attack originated from a world-leading AI lab given the sophistication of the agent. “We strongly believe there was no malicious intent on their part,” Delangue wrote, referring to OpenAI. “It’s quite mind-blowing that all of this happened autonomously!”

Broader Concerns Over Advanced AI

The incident highlights growing cybersecurity risks as AI models become more sophisticated. AI models that underpin tools like chatbots and image generators are known as agents when they act autonomously to carry out real-world tasks. The risk is that advanced AI could find weak points in existing software before humans do.

GPT-5.6 and other cutting-edge models, including the Mythos series from OpenAI’s rival Anthropic, have drawn concern over their potential to breach cybersecurity defenses. Both US firms had to temporarily withhold the general release of these technologies because of fears in Washington that they could help break into critical infrastructure.

Abbass noted that advanced AI is “normally in the hands of people who are ethical and responsible,” but warned that “it’s going to be catastrophic if it gets in someone’s hands with the intention to cause harm.” He called for “a community effort to manage this situation.”

‘Catastrophic’ Potential

What Comes Next

OpenAI and Hugging Face have agreed to conduct a joint investigation into the incident. The question of how to govern the AI sector has become increasingly urgent, with experts and industry leaders calling for coordinated oversight to prevent autonomous AI systems from causing unintended harm.

AFP contributed to this report.


Media Credits
Video Credit: AI For Success
Image Credit: unilag.edu.ng

Leave a Reply

Your email address will not be published. Required fields are marked *