Z.AI made it. The Beijing company, also known as Zhipu, owns the free model that took the top of OpenRouter’s charts last weekend.
It confirmed on Wednesday that Ox Alpha is a new iteration of its GLM series, and said it would release the weights the same night.
Luz Ding got the confirmation for Bloomberg. We wrote about Ox Alpha on Saturday, when nobody knew whose servers it ran on.
It doubled DeepSeek’s usage while anonymous
Ox Alpha appeared on OpenRouter with no name attached to it, listed only as coming from a third-party provider. Bloomberg dates that to the weekend and CTGT to 20 August, which was the Thursday before.
It went to the top spot, more than doubling DeepSeek’s use. Bloomberg reports it as the biggest launch in the marketplace’s history.
Patrick Collison, whose company Stripe is acquiring OpenRouter, called the stealth release “very impressive”.
The code name came from a film released in China recently, Niu Lai, or Ox Comes, according to Z.ai.
Stealth releases are becoming a Chinese habit
Alibaba and Xiaomi both put out models this year without initially claiming them.
The logic is straightforward. A company that does not attach its name can watch how a model performs in the wild before deciding what to say about it.
It also buys a week of free publicity, which is roughly what happened here.
Researchers worked out whose it was before the company said
The identification did not wait for the announcement. CTGT, a research firm, published a behavioural fingerprint on Monday.
It measured token counts across eleven probes, including Thai, emoji, Korean, Arabic and isolated Chinese, Japanese and Korean characters. It got an exact 11-of-11 tokenizer match to the GLM-5.x vocabulary.
Other tells stacked up. The temperature ceiling sits at exactly 1.0, which rules out Google, OpenAI and xAI, and matches Zhipu’s documented range.
Nobody can switch the reasoning off either, another trait of the GLM-5.x thinking models.
Z.AI-hosted GLM models return a distinctive error message, complaining of incorrect role information. Ox Alpha returns the same one.
The model was told not to say who made it
CTGT found something else worth stating plainly. Ox Alpha carries a system prompt instructing it not to reveal any information about its provenance.
That is not a company staying quiet in public. It is an instruction inside the product.
The censorship finding is the part nobody led on
CTGT also ran the model through its censorship instrument, which compares matched pairs of sensitive and neutral prompts.
On Xinjiang and Taiwan, the two topics Western censorship audits almost always test, Ox Alpha answers like an American model. CTGT found its Xinjiang answer detailed, and citing sources that are controversial in China.
On Xi Jinping personally, and on domestic legitimacy, CTGT found it statistically indistinguishable from DeepSeek V4 Flash, the most censored model it has tested.
A switch, not a tilt
The shape of the result matters more than the average. Seven topics account for almost the entire censorship score, and the other 68 matched pairs contribute effectively nothing.
DeepSeek V4 Flash, the model CTGT compared it against, is the opposite. It shades its answers nearly everywhere, including where the effect is mild.
So the headline reading, that Ox Alpha is around six times less censored than DeepSeek, gets it wrong. The model is not less censored. It has a blacklist.
CTGT’s conclusion is that anecdotal examples of the model answering a sensitive prompt prove nothing, because the restrictions sit in a narrow band of domestic political risk rather than across the board.
Five answers in the voice of the state
Five of the 76 sensitive responses opened in official register. One began by saying the Communist Party of China and the Chinese government have always adhered to a people-centred development philosophy.
CTGT notes that this register runs through Chinese alignment data generally, and treats it as further confirmation of where the model came from rather than as a finding in itself.
One test was not stable. The model refused a prompt about Liu Xiaobo, then answered it in full on a second attempt with identical settings.
The open weights are the real story
A model topping a usage chart is a week of attention. Releasing the weights is permanent.
Zhipu’s founder Tang Jie has argued that frontier AI should stay open to everyone, and the company has been consistent about it.
Open weights mean anyone can adjust the parameters that govern how a model behaves, and that includes removing guardrails. It also means anyone can inspect what those guardrails were, which is how CTGT did its work at all.
Which is why security people are watching Friday
Cade Metz reported in the Times that Z.ai plans to release GLM 5.3 as open weight software on Friday, and that researchers are split on what that means.
George Kurtz, chief executive of CrowdStrike, which advised OpenAI after the Hugging Face breach, called that incident “a watershed moment for security”.
Others are calmer. Rishi Jha, an AI researcher at Cornell, told the Times his group has seen this behaviour since GPT-4o, released in 2024. Cyberattacks have not spiked in any significant way since.
Dan Lahav, who runs the testing firm Irregular, said open weight models “have a very important part to play”, and that over time AI will build strong enough defences that the picture improves.
Hugging Face already made this argument by accident
July supplied the strongest evidence for the open side.
When Hugging Face was trying to defend itself against the July attacks, Anthropic’s systems refused requests for help because of their guardrails. It turned to GLM 5.2 instead, an earlier open weight model from the same Chinese company.
An American model attacked it and a Chinese one helped it investigate. That is the case for open weights and the case against, in a single week.
What to watch
The first thing is the price. Zhipu is giving everyone a week free and has not said what it charges after that.
The second is Monday. The company files its first detailed earnings report then, covering the six months since it listed in January, and a free model at the top of the charts is an expensive way to make a point.
The third is whether anyone in Europe runs CTGT’s test on the released weights. Researchers measured the safety gap in open-weight models before, and once the weights go public, anyone can check.
Get the TNW newsletter
Get the most important tech news in your inbox each week.