Export controls were supposed to settle the question of Chinese AI compute. Instead they produced a specific engineering answer: if you cannot buy the fastest chip, build the largest machine.
Huawei's answer is the CloudMatrix 384 supernode, marketed as the Atlas 900 A3 SuperPoD. Announced on 10 April 2025 at Huawei Cloud's ecosystem conference and put into production at the Wuhu data centre, it connects 384 Ascend 910C accelerators and 192 Kunpeng CPUs across 16 racks — twelve compute, four holding second-tier Unified Bus switches — into what the scheduler treats as one computer.
Key takeaways
- Scale-up domain: 384 Ascend 910C NPUs plus 192 Kunpeng CPUs in a 16-rack supernode, presented to software as a single machine.
- Compute: about 300 PFLOPS of dense BF16, which independent analysis (SemiAnalysis, April 2025) puts at close to double NVIDIA's GB200 NVL72, with roughly 3.6 times the aggregate memory capacity and 2.1 times the memory bandwidth.
- Interconnect: 3,168 optical fibres and 6,912 400G LPO modules forming the Unified Bus plane; cross-node bandwidth degradation under 3% relative to intra-node traffic, which is what makes single-machine scheduling possible.
- Power cost: the same analysis estimates about 4.1 times the power draw of the NVIDIA configuration, and roughly 2.5 times worse performance per watt.
- Deployment: at the World Artificial Intelligence Conference in Shanghai in July 2026, Huawei said the 384-chip supernode had been deployed in more than 750 commercial projects across internet, telecom and finance.
The trade, stated honestly
The design buys performance with electricity. That is not a secret and Chinese engineers discuss it openly: Huawei's rotating chairman Xu Zhijun described the breakthrough in May 2026 as the result of systematically driving down latency across chip, interconnect and system layers.
The point is that the trade is rational in its context. China's constraint is not primarily electricity — it is access to leading-edge silicon. A system that is 2.5× less efficient per watt but 2× more capable per rack, and buildable entirely from domestic parts, is a rational choice for a buyer who cannot buy the alternative at any price.
Analysts have noted exactly this: power is a weaker constraint in China than in the United States, which changes how the arithmetic looks to a Chinese operator.
The harder problem was never the chip
Compute hardware is the visible half. The reason NVIDIA still dominates globally is CUDA — two decades of operators, libraries, debugging tools and developer muscle memory.
Huawei's 2025 response was to give the stack away. At an Ascend computing summit in Beijing on 6 August 2025, the company announced a full open-source release of CANN (Compute Architecture for Neural Networks), the layer between AI frameworks and Ascend silicon. At HUAWEI CONNECT the following month, Ascend Computing committed to publishing all CANN operators on GitCode by late September, with domain libraries, the graph engine, Ascend C and the MindIE inference engine following in December, plus 1,500 PFLOPS of compute and 30,000 development boards a year for ecosystem work.
By WAIC 2026, Huawei reported the CANN community at 67 projects and more than 3,500 monthly active developers. That is small compared with CUDA. It is also the first time the number has been worth reporting.
MindSpore, Huawei's own framework, has been open source under Apache 2.0 since March 2020 and runs on Ascend, GPU and CPU backends.
What the supernode proves
The performance figures that matter are the serving ones. In a June 2025 paper with inference provider SiliconFlow, Huawei documented the system serving DeepSeek-R1 with 320-way expert parallelism and INT8 quantisation: prefill throughput of 6,688 tokens per second per NPU and decode throughput of 1,943 tokens per second per NPU within a 50 ms per-output-token budget.
That is the number that answers the export-control question. Not "can China build a fast chip" — but "can China serve frontier models on infrastructure it fully controls." The answer, in production and across 750 commercial projects, is yes.
The gap in per-watt efficiency and software maturity remains real. But it is now a gap on a curve, not a wall.
Specifications and deployment figures as disclosed by Huawei and reported at WAIC 2026; comparative performance per SemiAnalysis, April 2025.
