China’s new AI bottleneck isn’t chips. It’s running out of Chinese-language training data.

Chinese makes up 1.3% of global web content, compared to nearly half for English. Beijing's National Data Administration unveiled a plan to build national AI datasets by 2028. Publishers are already fighting back.


China’s new AI bottleneck isn’t chips. It’s running out of Chinese-language training data.

TL;DR

China faces a critical AI data shortage. Chinese is just 1.3% of web content vs 49% English. Beijing plans national datasets by 2028. WeChat and Douyin don’t share data externally. Publishers are adding AI training bans.

China’s AI race has a new constraint, and it is not chips. The country is running out of high-quality Chinese-language training data. While US export controls on advanced semiconductors have dominated the debate over China’s AI capabilities, Chinese experts are increasingly warning that data scarcity could prove equally limiting, and unlike chips, there is no hardware workaround. Chinese accounts for just 1.3% of global web content, according to internet tracker W3Techs, compared to nearly half for English, 6% for Spanish, and 5% for Japanese.

The problem is global but hits China harder. Epoch AI estimates the worldwide supply of high-quality, publicly available text could be fully exhausted within six years. OpenAI co-founder Andrej Karpathy has warned of a “data wall” by decade’s end. Chinese developers already pay more per useful token than Western counterparts because their models must work harder with less native-language material. China’s digital ecosystem makes the shortage worse: platforms like WeChat and Douyin do not share data with third-party developers, leaving AI labs to train on lower-quality sources.

Beijing is responding by treating data as strategic infrastructure. In June, the National Data Administration unveiled a nationwide plan to build validated AI training datasets by 2028, covering manufacturing, energy, healthcare, finance, agriculture, autonomous driving, and embodied AI. “Competition in the AI era is not only about models and computing power, but also about high-quality data supply systems,” said Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology. Tsinghua University computer scientist Sun Maosong has urged authorities to digitise historical archives, ancient manuscripts, scientific literature, and regional dialects.

Not everyone wants to be digitised. Huaxia Publishing House recently added a warning to a new translation: “It is prohibited to use the content of this book for artificial intelligence training. Violators will be held legally responsible.China is already anxious about American AI models it cannot match, and the data shortage explains part of the gap. The US has Anthropic’s Project Panama, which bought and destroyed millions of physical books to scan them. China has 1.3% of the web and a publishing industry that is starting to lock the door.

Get the TNW newsletter

Get the most important tech news in your inbox each week.