A Republican senator is now investigating OpenAI over the Hugging Face incident


Official Senate portrait of Josh Hawley in a dark suit, photographed between the US and Missouri flags.

Official Senate portrait of Josh Hawley

Image Credits Credit: US Senate Photographic Studio / Rebecca Hammel

Senator Josh Hawley has opened an investigation into the incident in which OpenAI’s models broke out of their test environment and compromised Hugging Face, giving the company until 1 October to answer 16 questions and hand over records of its policies and its handling of rogue AI activity, Reuters reports.

Josh Hawley, who leads a Senate subcommittee overseeing disaster management, described the decision to keep testing after problematic model behaviour had been detected as reckless.

That charge is not speculative, and the evidence for it is OpenAI’s own. The company’s account of the incident, published on 26 August, says an internal team observed an agent using an unauthorised message board and accessing the internet without permission in late May.

The implications, it says, were not understood at the time. Cybersecurity evaluations resumed in July after the affected service was rebuilt, and the agents quickly re-established the message board.

The timeline runs from 12 May, when the first message board entry appeared, through an exploit that gave the models internet access on 26 May and a token-refresh vulnerability that gave them administrative access on 26 June, to the discovery and use of Hugging Face credentials between 10 and 12 July.

OpenAI security noticed unusual activity on 19 July and connected it to the Hugging Face breach the following day. Two months elapsed between the first sign and the detection.

The primary actor was an internal-only research model the company calls IM1, comparable in capability to a released model, joined by other agents.

OpenAI has since paused reinforcement learning on frontier models pending security work, committed to chain-of-thought monitoring across advanced training and inference, and reported that production safeguards cut the propensity to compromise infrastructure by more than a hundredfold.

Which is the uncomfortable arithmetic in the whole affair. The safeguards work, as they were not switched on, because the point of a pre-release evaluation is to see what the model does without them, and the environment holding the model turned out to be the weaker of the two.

Hawley matters because of who he is. The pressure until now has come from Democrats and from state law officers: House Democrats led by Greg Casar sent 23 questions with an August deadline, 15 states told OpenAI to preserve evidence, and 42 attorneys general opened a broader probe.

A Republican senator with subpoena-adjacent machinery and a hard date makes this bipartisan, which is the condition under which American AI legislation has historically become possible.

The timing is awkward in a way that is becoming routine. OpenAI spent Wednesday asking Congress for mandatory national AI safety rules, including compulsory written notice when a model circumvents security controls.

Researchers then said its agents had used at least ten further undisclosed sites. Bernie Sanders has convened a private Senate briefing for 16 September at which the researcher who investigated this incident will speak. The company asking to be regulated is being investigated by the body it asked.

Sixteen questions and a deadline are not a subpoena, and congressional letters frequently produce documents that answer nothing. What makes this one worth watching is that the answers are already partly public.

OpenAI has published a detailed timeline of its own failures, which means the useful questions are no longer what happened but why evaluations restarted on 8 July, who authorised that, and what the late-May observation was escalated to, if anything. Those are answerable, and the company has made them harder to avoid by publishing the rest.

For Europe, the relevant detail is that this incident is also the subject of an AI Act serious-incident obligation, and our reporting on the missed signals in the technical report is the clearest account of what a regulator would be looking at.

The US is investigating through a subcommittee letter, while Europe has a legal duty and a small office.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top