ByteDance is preparing a real-time spatial video model under Zhang Yiming, Bloomberg reports

Built on Seedance and aimed at Pico headsets, the model would render interactive 3D worlds in the cloud at roughly 50 milliseconds of latency.


Close up ByteDance company logo on office building.

Close up ByteDance company logo on office building.

Image Credits Credit: Tang Yan Song via Shutterstock.com

ByteDance is preparing an AI model for real-time spatial video generation, with founder Zhang Yiming personally overseeing the work and a launch possible as soon as next month, Bloomberg reported, citing people familiar with the matter who asked not to be identified.

One of those people cautioned that the timing is not settled and the plans could change. TNW has not independently verified the account, and a ByteDance spokesperson did not respond to Bloomberg’s request for comment.

The model would be built on Seedance, ByteDance’s existing video generation system, and would let users create interactive virtual worlds for live streams, short-form dramas and games.

The reported specification is the interesting part: on-demand video at around 20 frames per second with latency of roughly 0.05 seconds, generated in the cloud rather than on the device.

That last detail is the strategy. Rendering spatial content remotely takes the computational load off the headset, which lowers what the hardware has to do and therefore what it has to cost.

ByteDance owns Pico, its extended reality arm, and the model is reportedly meant to generate worlds that respond to Pico users’ voices and movements.

If it works, the contest over XR moves away from hardware specifications and towards models, cloud capacity and content distribution, three things ByteDance already has.

None of this arrived from nowhere. 36Kr reported earlier this year that ByteDance had set four AI priorities for 2026, with world models at the top of the list, ahead of holding Seedance’s lead in video, improving coding, and commercialising Doubao.

World models were said to command the company’s largest data budget of any model direction, an eight-figure sum in renminbi that 36Kr’s sources put at three to four times what rivals were spending.

The same reporting set the target explicitly: ship at least one world model by the end of the year and measure it against Google’s Genie, which now lets users walk around Street View imagery rendered in real time.

Internal testing early in 2026 put ByteDance about 10% behind the global state of the art, on the same account. A launch next month would be ahead of that schedule.

The company is pursuing two routes at once, according to 36Kr: a vision-language-action approach aimed at embodied intelligence and robotics, and 3D simulation for entertainment and games.

The spatial video model belongs to the second, which is also the one with an existing user base attached to it.

A world model, in the sense everyone is now using, is a system that learns how environments behave well enough to render one that responds coherently to what a user does in it.

The reason video companies keep turning up in this field is that the training material is video, and the firms with the most of it, and the most experience compressing it into something that renders fast, start from an unusual position. ByteDance has spent a decade building exactly that pipeline for a different purpose.

Bloomberg frames Zhang’s involvement as putting him among researchers such as Fei-Fei Li and Yann LeCun, who have argued that models grounded in visual and physical understanding, rather than language alone, are the route to systems that can act in the world.

That case is now being tested commercially by companies whose actual product is entertainment.

The money behind it is not in doubt. ByteDance secured a $30bn loan last week, Bloomberg reported, and has been weighing capital expenditure of as much as $70bn on its AI build-out.

Seedance already underpins CapCut and Doubao, and the company remains, on its own domestic terms, a challenger against Alibaba, DeepSeek and Moonshot AI.

For Meta and Apple, the implication is awkward rather than immediate. Both have spent heavily on headsets that have not gone mainstream, and Europe’s own XR specialists have built their businesses on high-end hardware rather than cheap devices fed from a data centre.

A cloud-rendered world model does not beat a Vision Pro on fidelity. It just makes the fidelity somebody else’s problem.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top