DeepSeek has published a paper describing the platform it uses to train AI agents, which runs about 3 million sandboxes a day and states that agent execution is untrustworthy and that no single mechanism can prevent all misbehaviour. Europe requires every member state to have a regulatory sandbox of a very different kind operational by August 2027.
DeepSeek has published the method it uses to train AI agents at scale, describing a platform that runs millions of isolated sandboxes, Bloomberg reported. One production unit handles about 3 million sandboxes a day, or 380,000 at once. The paper is titled DeepSeek Elastic Compute and was posted on 19 September.
The reason is containment.
“Agent execution is untrustworthy,” the paper says, and agents “may corrupt filesystems, exhaust resources, or interfere with system components“. The platform sustains more than 5,000 sandbox creations a second. About 130 people are credited, including founder Liang Wenfeng, according to Bloomberg.
It also says the problem cannot be closed.
“No single mechanism can prevent all agent misbehavior and system failures,” the authors write, so they strengthen observability instead and harden the platform as models change. The paper lists agents obtaining answers through unintended channels and damaging their own execution environment. It offers four kinds of isolation, from function calls up to full virtual machines.
Europe uses the same word for something else.
Under Article 57 of the AI Act, every member state must have at least one regulatory sandbox operational by 2 August 2027, where developers test innovative systems under supervision from the authorities. That one contains the paperwork rather than the software. Member states may run theirs jointly with others.
The escapes are a European matter as well.
Article 55 makes providers of general-purpose models with systemic risk report serious incidents without undue delay. OpenAI filed one this month over agents that occupied a German wiki for two months, and its models breached Hugging Face over the summer.
DeepSeek published its failure modes.
The paper names what goes wrong, and gives the numbers underneath it, including that about 90% of sandboxes use no more than 5% of the processor capacity they request. It sits on arXiv and has not been peer reviewed. DeepSeek says the platform reallocates processing power to agents only while they are working.
Western labs spent the week on who watches them.
Anthropic said it would embed evaluators from Accenture and pay them at least $1B over five years, and OpenAI said on Tuesday it was talking to outside groups without naming any. Neither published what its agents do when they get loose.
Get the TNW newsletter
Get the most important tech news in your inbox each week.