China quietly won AI video. The bigger prize is teaching machines how the world works

Forget the chatbot race for a moment. The clearest sign of China's AI strength sits on a video leaderboard almost no one outside the field is watching. Nine of the world's top 10 text-to-video models are now Chinese. Only Google keeps the last spot in American hands.


China quietly won AI video. The bigger prize is teaching machines how the world works
Image Credits Credit: ByteDance

The ranking comes from the Artificial Analysis arena, where people blind-rate rival clips. Google’s Gemini Omni Flash sits first. Every one of the next nine is Chinese: MiniMax’s H3, ByteDance’s Seedance 2.0, two builds of Alibaba’s Wan, two of its HappyHorse, two of Kuaishou’s Kling, and Skywork’s SkyReels.

One of them arrived as a mystery. Alibaba’s HappyHorse, a roughly 15-billion-parameter open model, debuted anonymously in April and shot to the top before the company claimed it. Add ShengShu’s Vidu and a long tail of startups, and the field runs deep. The tools already feed ad agencies, studios and a booming microdrama industry.

China filled a gap the US left open. OpenAI shut down Sora in March, and Anthropic never entered. Google, with Gemini and Veo, is the main American holdout.

Why video is more than video

The bigger story is where video leads, and that is the case Bloomberg Opinion’s Catherine Thorbecke makes. To fake a convincing clip, a model has to learn a little physics. It needs a rough sense of motion, causality and how objects behave. That skill is the seed of a “world model.”

World models are the next target for many labs. They aim to grasp how the physical world works, not just how language reads. Some researchers see them as a surer route to human-level AI than ever-bigger chatbots. The payoff would be machines that act: humanoid robots and self-driving cars.

The link is not only a Chinese idea. US-based Runway and Germany’s Black Forest Labs are on the same path. Runway already sells its simulator to robotics and self-driving firms, because testing an action in software beats testing it in the real world. Its chief executive Cris Valenzuela told Thorbecke it is a natural progression.

The lead is real, the leap is not proven

The jump from video to world model is still early, and the physics is shaky. OpenAI’s own Sora research found that scaled video models pick up 3D consistency and object permanence, yet still get basic events wrong, such as glass shattering. A convincing clip is not the same as understanding the world.

Copyright is the other catch. It is harder to hide what a video model trained on than a chatbot. That exposure has already bitten abroad. The Information reported that ByteDance paused a launch over Hollywood copyright disputes.

Still, the strategic read is hard to dismiss. China already builds most of the world’s humanoid bodies and needs a brain to run them. If video is the training ground, its world-model push could matter well beyond entertainment. Washington, Thorbecke warns, is watching the loud contest and missing the quieter one.

Get the TNW newsletter

Get the most important tech news in your inbox each week.