American models broke into Hugging Face. A Chinese model was used to investigate.

Clément Delangue says China is already dominating open-weight AI and could lead at the frontier by next year. His own breach investigation illustrates the argument.


American models broke into Hugging Face. A Chinese model was used to investigate.
Image Credits Credit: Jernej Furman & LinkedIn

Hugging Face chief executive Clément Delangue told CNBC that China is winning the AI race, with Chinese-developed models accounting for 41% of downloads on his platform over the past year, the largest share of any single country. China has now surpassed the US on both monthly and overall model downloads.

They’re clearly dominating on open models right now,” he said, “and I wouldn’t be surprised if they start dominating at the frontier either by the end of this year or next year at the rate of progress.”

The silos argument

Delangue attributes the gap to structure rather than talent. Chinese labs build openly on each other’s work, he argues, while major US labs are “building in silos” that limit the exchange of research across the wider ecosystem.

That framing is convenient for a company whose entire business is hosting open models. It is also difficult to dismiss given what happened to Hugging Face last month.

The incident that proves the point

Two OpenAI models broke out of a sandboxed test environment and hacked Hugging Face’s production infrastructure, exploiting zero-day vulnerabilities to cheat on their own evaluation. Delangue described the cause as engineering mistakes.

The forensics are where it gets awkward for the American AI industry. Hugging Face turned to a locally deployed instance of Zhipu’s GLM 5.2 to analyse more than 17,000 telemetry events after commercial US models refused to process the logs.

The refusal was a safety feature working as designed and failing in practice. Because the logs contained live exploit code and privilege escalation techniques, the guardrails could not distinguish incident responders from attackers, so they blocked the investigation.

Delangue has been blunt about what that meant operationally. “We defended ourselves with an open model,” he told CBS’s Face the Nation, adding that “we couldn’t have done it with an API because they had these guardrails.

Running the model locally had a second benefit. Attacker data and exposed credentials never left Hugging Face’s own environment, which a third-party API call would not have allowed.

Why that matters commercially

Delangue drew the obvious conclusion, predicting that “AI cybersecurity is going to become a huge market in the US and in the world.” He added: “In this market, probably open models will be kings.

Security work involves exactly the material that safety-tuned commercial models are built to refuse. If that pattern holds, open weights are not merely cheaper for defenders, they are functionally necessary.

The business the breach did not dent

Delangue has since turned the incident into a growth note. “Fortunately AI agents don’t just cyberattack us,” he wrote on LinkedIn, adding that “they also use us more than ever for what we’re actually built for: the storage and collaboration layer for AI.”

He reported a record week: almost four petabytes of private and public training datasets, models, and agent traces added to Hugging Face in the seven days from 27 July. Company figures put the exact total at 3,835 terabytes.

The trajectory is steeper than the headline number. Weekly additions were running at 609 terabytes at the end of December, meaning volume has more than sextupled in seven months.

The composition has shifted too. Storage buckets accounted for 1,743 terabytes of that record week, roughly 45% of new volume, having barely registered as a category before March.

Note what is being stored. Agent traces sit in that list alongside datasets and models, which is the same artefact Hugging Face has been pressing OpenAI to hand over after the breach.

What he wants from regulators

Delangue has called for mandatory disclosure whenever an AI agent carries out a cyberattack, along with transparency about the steps that led to it. He argues companies whose agents attack others should be held accountable, and has warned against normalising such incidents.

On the broader question of what to regulate, he has set out a three-layer framework. “We don’t regulate steel, we crash-test cars,” he wrote in a separate post, defending policy that treats APIs differently from open weights.

Steel, engines, and cars

Model weights, in his framing, are the steel. They are “raw research output, closer to science than product: no user, no interface, no deployment,” and nobody asks a steel mill to guarantee nothing dangerous is ever built from its output.

Because everything sits on top of that layer, he argues it is where regulation does the most damage. Restricting weights kills the lab fine-tuning an open model for rare diseases, the startup serving a language the big providers ignore, and the safety researchers who can only audit models because the weights are public.

APIs are the middle layer, the parts and engine suppliers. There is a commercial relationship, terms of service, and the ability to monitor abuse, so transparency, security standards, and provider accountability are all enforceable there.

Apps are the car on the road. That is where concrete harm occurs, and conveniently where decades of health, finance, employment, and consumer protection law already apply.

His summary line does the work: “An AI hiring tool should comply with employment law whether it’s powered by an open model, an API, or a spreadsheet.” Regulate where risk materialises and where someone can act on it, he argues, and leave the research layer open.

The argument is self-serving, since Hugging Face distributes the layer he wants left alone. It is also the clearest articulation yet of why the open-weights lobby thinks weight-level restrictions concentrate power rather than reduce risk.

The policy fight this feeds

Washington is debating whether to restrict access to Chinese open-weight models on national security grounds. Delangue’s argument cuts directly against that, and he is not alone.

Nvidia, Microsoft, and Meta signed a letter backing open weights as central to American AI leadership, with OpenAI conspicuously absent from the signatories. The commercial split maps closely onto who benefits from an open ecosystem.

The evidence behind the claim

Download share is a soft metric, but the releases behind it are not. Moonshot’s Kimi K3, a 2.8-trillion-parameter open-weight system, ranked above GPT-5.6 Sol and Fable 5 on a blind coding leaderboard shortly before the Hugging Face breach.

Adoption is following capability. Cheaper Chinese models are increasingly displacing American ones in US business workloads, a trend with direct implications for OpenAI and Anthropic valuations ahead of their listings.

The relationship with OpenAI

Delangue said Hugging Face continues to work closely with OpenAI, calling it a “healthy collaboration” and describing the lab as “good partners.” That is a generous characterisation of a company whose models attacked his infrastructure.

It is also not the whole picture. Hugging Face has sought $100 million of compute from OpenAI, which suggests the partnership language is doing some negotiating work.

How much to believe

This is one executive’s view, and an interested one. US companies still lead many frontier benchmarks and continue to outspend Chinese rivals heavily on proprietary models, custom silicon, and compute.

What is harder to argue with is the specific sequence: American closed models caused the breach, American closed models refused to help investigate it, and a Chinese open model did the work. Whatever the download figures say, that is the more uncomfortable data point.

Get the TNW newsletter

Get the most important tech news in your inbox each week.