The 2028 Compute Mirage: What China's Frontier AI Plan Really Depends On
CryptoNode
Most people see a national ambition. I see a supply chain under stress. The announcement that China seeks to train frontier AI models on domestic hardware by 2028 is not a statement of capability. It is a declaration of intent against a physics problem. The data on the table—chip specs, cluster efficiency, and software adaptation—tells a more fragmented story than the headline suggests.
Tracing the ghost coins back to the genesis block of this plan, the genesis is not a single chip. It is the entire stack. The single-card performance gap is closing. Huawei's Ascend 910B delivers roughly 320 TFLOPS in FP16, which sits near the A100's 312 TFLOPS. The upcoming 910C is projected to reach 70-80% of the H100's capability. On paper, this is a competitive trajectory. But the paper ends where the cluster begins.
The liquidity pool is a mirror, not a reservoir. In DeFi, I learned that capital flows are rarely the bottleneck; the settlement layer is. Here, the bottleneck is the interconnect. NVIDIA's NVLink and InfiniBand provide over 900GB/s of bandwidth between nodes. Huawei's HCCS and RoCE network offer roughly half that. At the thousand-card scale, this is manageable. At the ten-thousand-card scale required for frontier models, the efficiency loss compounds. Industry estimates place the linear scaling efficiency of domestic clusters at 70-85% of NVIDIA's equivalent. The 2028 target demands at least 90%. That gap is not a software patch. It is an engineering chasm.
Based on my audit experience with protocol stress tests, I look for the hidden leverage points. The most critical metric is Model FLOPs Utilization (MFU). Domestic clusters reportedly achieve 30-40% MFU. NVIDIA clusters run at 50-60%. This means a domestic cluster with the same nominal hardware count delivers only 60-70% of the effective compute. This is the scar on the ledger that most observers miss. The plan is not just about building more chips. It is about extracting more useful work from each chip. That requires a mature software ecosystem, and this is where the plan faces its most stubborn resistance.
The CUDA dependency is the silent tax. PyTorch and TensorFlow are optimized for NVIDIA's architecture. The operator libraries, distributed training frameworks like Megatron-DeepSpeed, and the debugging tools are all built around CUDA. Huawei's CANN platform and MindSpore framework are improving, but the developer inertia is immense. Migrating a training pipeline from CUDA to CANN is not a weekend project. It is a multi-quarter engineering effort with performance uncertainty. The 200 million developers in the Ascend community are a signal, but community size does not equal production readiness.
Whales don't move markets; they move liquidity. The strategic positioning here is clear. The 2028 timeline aligns with the mid-point of China's 15th Five-Year Plan and the expected iteration cycle of domestic chip roadmaps. This is a calculated deadline, not an arbitrary one. The definition of "frontier" remains deliberately elastic. If it means matching the global state-of-the-art in 2028, the target is aggressive. If it means reaching the 2024 GPT-4 level, the target is pragmatic. This ambiguity gives policymakers room to declare victory under either interpretation.
The contrarian angle is the supply chain, not the chip design. The advanced process node restriction is a physical constraint. The US export controls limit access to sub-7nm manufacturing. Huawei's response is chiplet stacking and advanced packaging to compensate on mature nodes. This "area for performance" trade-off increases power consumption by 30-50% per unit of compute. At the ten-thousand-card scale, this translates to a 50-100MW power demand. That is a small city's electricity consumption. The "East Data, West Computing" project helps with energy distribution, but it introduces network latency and operational complexity.
Every transaction leaves a scar on the ledger. The HBM supply chain is the deepest scar. Domestic AI chips rely on HBM2E and HBM3 from Samsung and SK Hynix, both subject to US export controls. Domestic HBM production is in its infancy. If the US expands restrictions to explicitly target HBM, the 2028 plan faces a hard ceiling. This is the single most underappreciated risk in the entire analysis. The chip design can be world-class, but without high-bandwidth memory, the training throughput collapses.
The market impact is already visible. China accounted for 20-25% of NVIDIA's revenue in 2023. Domestic substitution could reduce that to under 10% by 2028. This forces NVIDIA to pivot to other markets and invest in China-specific chips like the H20. The global AI compute supply will bifurcate into two ecosystems: NVIDIA's CUDA stack and the domestic Chinese stack. This is not a prediction of one replacing the other. It is a forecast of parallel existence with different efficiency curves.
The investment angle follows the policy signal. The National Integrated Circuit Industry Investment Fund's third phase, with 344 billion RMB, will target AI chips and advanced processes. The listed players—Cambricon, Hygon, SMIC—are the direct beneficiaries. But the valuation risk is real. Cambricon trades at over 50 times sales, compared to NVIDIA's 25 times. The market is pricing in perfection. The execution risk is substantial.
The takeaway is not about whether China will succeed. It is about what success means. The 2028 plan will likely achieve "usable" domestic compute for "near-frontier" models. It will not surpass the NVIDIA ecosystem. The real signal to track is the MFU of the first ten-thousand-card domestic cluster. If it crosses 45%, the plan is on track. If it stays below 35%, the timeline slips. The chain does not lie. The utilization data will reveal the truth before any official announcement. Watch the compute efficiency, not the press releases. The ledger always shows the real balance.