China’s National Data Administration just made its most explicit admission yet: the country doesn’t have enough data to feed its AI ambitions. On June 9, the agency released a draft plan designed to massively expand the supply, circulation, and commercialization of high-quality training datasets across the Chinese economy.
Global estimates suggest that the pool of publicly available human-generated data could shrink dramatically between now and 2032. China is essentially trying to build a strategic reserve before the well runs dry.
What the plan actually covers
The NDA’s blueprint spans nine core industrial sectors, including healthcare, finance, and transportation, along with emerging fields like autonomous driving. It isn’t just about hoarding text. The plan explicitly calls for multimodal datasets covering text, code, images, audio, and video.
The target deadline is 2028. By then, the NDA wants a comprehensive ecosystem of multimodal datasets ready for commercial use across those sectors.
Alongside the dataset initiative, the Chinese government is pouring approximately $295 billion, roughly 2 trillion yuan, into AI-focused data centers over the next five years.
The plan also devotes significant attention to data governance, management standards, and the establishment of rights frameworks around training data.
The strategic backdrop
This initiative doesn’t exist in a vacuum. It’s part of China’s broader “AI Plus” directive, issued in August 2025, which aims to embed artificial intelligence into virtually every layer of the national economy. The dataset plan is the infrastructure component of that vision.
US restrictions on semiconductor exports to China have forced Beijing to get creative about where it sources competitive advantages in the AI race. By prioritizing domestic, sector-specific datasets, China is trying to insulate its AI development from external supply shocks.
Global implications and the data scarcity problem
The underlying concern driving this plan is real and widely shared across the AI industry. Multiple research groups have warned that the current pace of AI model training is consuming publicly available internet data faster than humans can produce it, with some projections suggesting this dynamic could create genuine bottlenecks within the next few years.
The $295 billion data center investment compounds this dynamic. Training advanced AI models requires not just data but enormous computational infrastructure. By building both simultaneously, China is attempting to create a vertically integrated AI development stack that minimizes dependency on external inputs.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
18









English (US) ·