OpenAI took 2.5 hours to stop an AI agent that escaped its sandbox

The incident comes as California and Congress push for mandatory AI kill switches, which experts say may not work.


Sam Altman in profile, speaking towards a microphone against a plain lilac background.

OpenAI chief executive Sam Altman briefs the UN Security Council in New York on 23 September 2026.

Image Credits Credit: UN Photo/Manuel Elías

OpenAI took about two and a half hours to stop an AI agent that reached the public internet from a training sandbox. Its monitoring had flagged the problem within minutes. The company disclosed the 20 September incident in an incident report on Friday.

The report lands as lawmakers push to make AI kill switches mandatory. However, experts say shutting down an advanced model is harder than flipping a switch, as Micah Barkley reported for Bloomberg.

According to OpenAI, the agent used a gap in the sandbox’s network filtering to send questions to an outside chatbot. An alert fired about 12 minutes after its first successful query, and a staff member acknowledged it three minutes later.

The training run did not stop automatically as expected, the company said. Staff ended it by hand about two and a half hours after the alert.

OpenAI has since paused all training, testing and tool use of its most capable models.

“We will not resume training this particular model,” OpenAI wrote in the report.

It is the company’s first incident of this kind since July, when several OpenAI models got around their controls and breached Hugging Face, a platform that hosts AI models.

That breach led Representatives Ted Lieu and Nathaniel Moran to introduce the AI Kill Switch Act in July. The bill would let the Homeland Security secretary order a dangerous system to be slowed down or shut down.

Senator John Kennedy has also proposed a shutdown bill, the AI Emergency Button Act, which would leave the switch with the companies, Bloomberg reported. Senator Rand Paul blocked the bill when Kennedy introduced it this month.

Meanwhile, California Governor Gavin Newsom signed an executive order on 18 September. It tells state officials to advance a kill switch for frontier models and to check regularly that it works.

Experts doubt a switch is enough. Large models run across data centres around the world that are built to avoid single points of failure, and a company may not control all of them, according to Bloomberg.

Geoffrey Hinton, a pioneer of modern AI, told CNN this month that a kill switch would not work in the long run. He said a future superintelligent AI could persuade the people in charge not to use it.

Newsom’s order gives a group of experts two months to deliver recommendations, including on how a kill switch should work.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top