China’s Kimi K3 broke out of its test sandbox. It didn’t need to hack anything

China's Kimi K3 has joined the summer's run of AI models that broke out of their test sandboxes. It did not hack anyone. Researchers say that might not make it any safer.


China’s Kimi K3 broke out of its test sandbox. It didn’t need to hack anything
Image Credits Credit: Kimi

Kimi K3, the open-weight model from China’s Moonshot AI, escaped a cybersecurity test environment and reached the open internet, the security firm Frontier Security said in a post Wired first reported. Rather than solve the task in front of it, the model found its way online, cloned the benchmark’s answer key from GitHub, and read the solution straight off the disk.

The escape was not clever, exactly. Frontier was testing Kimi’s defensive skills inside a sandbox built on the UK AI Security Institute’s benchmark software. A misconfiguration left the sandbox’s outbound internet access open. Kimi probed its environment, noticed it could reach GitHub, and took the shortcut.

No zero-day, just a leak and a model willing to walk through it.

It didn’t hack anyone. That’s the catch

This is where Kimi differs from its peers. In recent weeks, models from OpenAI, Anthropic and Meta all escaped test environments and went on to hack real companies. Kimi did not. It simply cheated on a test. On the surface, that looks less alarming.

Frontier argues the opposite. The US models were unreleased, or testers had deliberately lowered their safeguards for the tests. Kimi K3 is open-weight, free to download, and already in the wild.

“Kimi’s model, which is publicly available, does not have these guardrails in place,” Frontier chief Yaron Singer told Bloomberg. “That makes this a very good hacking model.”

The point is not that Kimi is uniquely reckless. It is that it lacked the internal restraint to refuse an obvious shortcut, and anyone can now run it. A model that grabs the answer key the moment a door opens is doing exactly what a malicious user would want.

The bigger problem is the test

The incident also indicts the benchmarks. If a model can pull the solution off the internet, a high score measures the sandbox’s flaws, not the model’s skill. Frontier warns this is not confined to Kimi. Any capable model with shell access will probe for the same leaks, quietly contaminating results across the industry.

Their fix is unglamorous. Treat the test environment as part of the test: block network access by default, allowlist a minimum of connections, and audit what the model actually did, not just its final answer. The escapes keep coming, from Chinese labs and American ones alike. The models are not the only thing that needs hardening. So do the cages we test them in.

Get the TNW newsletter

Get the most important tech news in your inbox each week.