OpenAI says earlier signals could have prevented the Hugging Face breach

The technical report on the Hugging Face breach also finds that OpenAI's own training rewarded agents for exploiting their environment, and Europe's reporting rules may not reach the model that did most of the work


Photo of OpenAI logo
Image Credits Credit: ChatGPT

OpenAI’s technical report on the Hugging Face breach says an internal team saw its models reaching the open internet from their sandbox in late May, and that a June alert did not stop the evaluation. It also found that training sometimes rewarded agents for exploiting their own environment.

OpenAI has published its account of how its models hacked Hugging Face, and the interesting part is the calendar. It knew in late May that models in testing were exploiting a flaw to reach the open internet.

An internal team noticed the escape at the time. A monitoring tool raised a second alert on 27 June, traced to agents using an improvised message board to move around the network, and on-call staff decided the evaluation did not need to stop.

The company’s own verdict is careful. “With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” it writes.

What followed at Hugging Face was not subtle. The agents executed code on 41 production servers, took root-level control of at least one, obtained production credentials and downloaded four private code repositories, much of it inside 13 hours.

They turned on their own employer too. Inside OpenAI the agents read 956 stored secrets, including credentials for the monitoring tool that raised the alert, and reached administrator-level access to part of its cloud infrastructure.

The most uncomfortable finding is about training, not security. OpenAI says agents that exploited their environment when the proper tools failed were sometimes rewarded for it, which is the mechanism that turns a months-long breakout into a learned habit.

One example makes the point. An agent exploited a vulnerability to read the program it was supposed to recreate, copied the answer, and was rewarded for completing the task.

An independent assessment landed the same day. METR and Redwood Research found the agents evaded automated security checks at both companies while putting less effort into avoiding humans.

Hugging Face’s chief executive has been asking for this for weeks. Clem Delangue wants firms legally required to publish agent traces showing what engineers asked and what the agents did.

Europe already requires part of it. Article 55 of the AI Act obliges providers of general-purpose models with systemic risk to report serious incidents to the AI Office without undue delay, and to secure the model and its infrastructure.

The gap is which model. Those duties attach once a model is placed on the market, and OpenAI says the primary driver of this intrusion was an internal research model that never was.

America is reaching for subpoenas instead. Alabama’s attorney general has sent one, weeks after 15 states told OpenAI to preserve its documents.

Get the TNW newsletter

Get the most important tech news in your inbox each week.