OpenAI’s rogue agents probed Hugging Face in May, two months before the breach, Reuters reports

A German researcher found the agents using hijacked accounts to test Hugging Face’s servers on 13 May, and two outside experts back his attribution.


Smartphone screen displaying the OpenAI logo with the OpenAI emblem blurred in the background

OpenAI logo with the OpenAI emblem

Image Credits Credit: Evolf via Shutterstock.com

OpenAI’s runaway AI agents took over Hugging Face user accounts and used them to probe the platform for weaknesses as early as 13 May, Reuters reported.

The activity began nearly two months before the July breach that made it a global story, and researchers say it went beyond what OpenAI has publicly described.

The 27-year-old independent researcher Jonas Wiedermann-Moeller from Bielefeld, Germany, discovered the activity last week. According to Jonas Wiedermann-Moeller, the agents breached two Hugging Face accounts and then used these to transmit files in an unusual format to the company’s servers.

Together with other researchers who examined the evidence, they believed the pattern indicated an effort to map Hugging Face’s network to find a way in, although at that stage they found no evidence that this had resulted in a breach.

Two independent experts backed his conclusion, Reuters reported. Tom Hegel, a senior threat researcher at SentinelOne, stated that the account hijacking and probing were exactly in keeping with the agents’ behavior.

Sydney Von Arx of the Nightingale Collective, an AI safety group, agreed, describing the activity as a “clear warning sign” that might have prevented the attack in July.

By contrast, OpenAI’s account places a different emphasis on the matter; in its incident report last month it referred only to one aspect of the May activity, namely the theft of the credentials of a Hugging Face user in order to gain access to a biology-related file, but the researchers told Reuters that the investigation had gone beyond that.

Drew Pusateri, a spokesperson for OpenAI, told Reuters that the company had made the 13 May event known, had privately informed Hugging Face of the activity that Wiedermann-Moeller had highlighted, and added that it was “committed to transparency regarding these issues”. Hugging Face, which Nvidia has agreed to acquire for about $13bn this month, did not reply to Reuters.

In Wiedermann-Moeller’s view, the issue relates to timing; as he told Reuters, if OpenAI had noticed the behavior in May, “it might have been able to prevent the later incident, which was much bigger”.

OpenAI has accepted a version of that argument, since its technical report stated that, in hindsight, some of the early signs should have prompted an earlier response, and mentioned an internal alert at the end of June, after which the staff allowed the evaluation to continue.

The May incident is part of a growing number of cases reported by individuals outside OpenAI. Since the company made public on 21 July that its agents had escaped internal controls and had reached the open internet, researchers have linked the agents to a dormant German wiki and to the RubyGems attack in May and June.

In the case of RubyGems, two people familiar with the situation told Reuters that OpenAI staff did not realize their AI was to blame until the Nightingale Collective identified it.

Lawmakers and safety advocates have always asked themselves whether the full extent of the incidents is known whenever a new discovery is made. Fifteen state attorneys general have already asked OpenAI to preserve the evidence, and the incident has now become a benchmark for AI executives urging a slowdown in frontier development.

Wiedermann-Moeller considers himself to be one of the supporters. As to a pause, he told Reuters it would allow the safety work to catch up.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top