Newsroom
Huawei

Huawei's Ascend Chips Have Overtaken Nvidia in China's AI Market

D
Debby Wang
September 28, 2026•12 min read•Updated September 28, 2026
Share:
Huawei's Ascend Chips Have Overtaken Nvidia in China's AI Market

Huawei's Ascend Chips Have Overtaken Nvidia in China's AI Market

TL;DR

Nvidia's share of China's high-end AI training chip market has effectively collapsed since October 2023 — not from a Huawei performance breakthrough, but from a stroke of U.S. export law. Huawei's Ascend 910B is now the default choice for new AI cluster deployments from Beijing to Shenzhen. The documented market gains are real; the capability parity claim is not.

Key Takeaways

  • The U.S. Bureau of Industry and Security's October 2023 rule update banned Nvidia's H800 and A800 from Chinese buyers, removing the only Nvidia chips legal for high-end AI training in China, according to the Bureau of Industry and Security.
  • Huawei reportedly shipped over 70,000 Ascend 910B units in 2023 alone, according to reporting by The Information; Huawei has not published official shipment data and the figure is unconfirmed.
  • ByteDance, Alibaba Cloud, and Baidu have each confirmed Ascend 910B use in AI training workloads through public infrastructure documentation and company disclosures across 2024.
  • Alibaba's Qwen team published Ascend 910B support for Qwen 2.5 model training and inference in its technical documentation — a first-party confirmation that the software stack is no longer prototype.
  • Canalys estimated Huawei at approximately 30–40% of China's AI accelerator market by revenue in 2024; treat this as a directional signal, not a precise figure, as analyst methodologies differ materially.
  • DeepSeek-V3's technical report explicitly documents training on 2,048 H800 GPUs stockpiled before the October 2023 export tightening — claims that DeepSeek's frontier training runs on Ascend hardware are unverified.
  • CITIC Securities projected Huawei could exceed 50% of China's new AI server deployments by end of 2025 — a projection, worth tracking against actual results.

The Rule That Made Huawei Dominant

In October 2022, the U.S. Commerce Department drew its first hard line: no A100 or H100 exports to China. Nvidia responded fast. The H800 and A800 were engineered specifically to stay just below the new performance threshold. Chinese hyperscalers ordered heavily. The workaround worked, briefly.

Then came October 2023. The Bureau of Industry and Security updated the rule again, closing the gap. The H800 was out. The A800 was out. There were no more legal Nvidia chips available for high-end AI training in China.

That is the precise moment Huawei became dominant — not because Ascend 910B crossed some performance rubicon, but because the competition was legislated out of the market.

It would be intellectually lazy to stop there. Huawei's hardware and software ecosystem was not ready to fill this gap in 2022. By 2024, the weight of the evidence suggests it largely is — with caveats worth naming precisely.

What the Ascend 910B Actually Is

Huawei's Ascend 910B is a second-generation AI accelerator from HiSilicon, Huawei's in-house chip design arm. It targets the training tier directly — the same workloads Nvidia's H800 was built for.

What is verified: ByteDance has publicly confirmed Ascend deployments for AI infrastructure. Alibaba Cloud's documentation lists Ascend 910B as a supported compute option for large model training. Baidu's AI Cloud platform has been built significantly on Ascend hardware. These are not anonymous supply-chain rumors. They are public infrastructure disclosures from companies with legal and reputational reasons to be accurate about what their clouds run.

What is unconfirmed: Performance comparison numbers circulating in Chinese tech media — claims that the 910B matches or exceeds H800 throughput on specific LLM training jobs — come largely from internal benchmarks that have not been independently replicated. The MFU (model FLOPs utilization) figures cited by Huawei and some Chinese model labs have not been verified by third-party analysis to the standard that MLPerf provides for Nvidia hardware.

The honest summary: Ascend 910B is production-capable for large-scale Chinese AI model training. Whether it approaches the efficiency of H100 — the chip Chinese companies cannot legally buy — remains genuinely unclear.

ChipVendorChina Legal StatusConfirmed at-Scale UsersIndependent Performance Benchmark
H100NvidiaBanned (Oct 2022)Pre-ban stockpiles onlyYes (MLPerf)
H800NvidiaBanned (Oct 2023)Pre-ban stockpiles; DeepSeek-V3 trainingPartial (Nvidia published specs)
A800NvidiaBanned (Oct 2023)Pre-ban stockpilesPartial
Ascend 910BHuaweiFreely availableByteDance, Alibaba Cloud, BaiduNone third-party verified
Ascend 910CHuaweiNot yet shipping (as of Q3 2026)N/AUnknown

The Software Gap Is the Real Story

Here is where the market-share figure becomes less interesting than what it does not tell you.

Nvidia's dominance in AI has never been purely about hardware. It is about CUDA — the programming layer, the libraries (cuDNN, cuBLAS), the toolchain that every ML engineer has spent a decade optimizing against. When you buy H100 clusters, you are buying into that accumulated ecosystem.

Huawei's equivalent is CANN — Compute Architecture for Neural Networks. CANN is functional. Major Chinese model labs have teams dedicated to it. But the gap between CANN and CUDA in developer tooling maturity, third-party library support, and debugging infrastructure is real. It is documented in engineering blogs from Alibaba, Baidu, and independent ML practitioners in China — not just in Western analyst speculation.

The practical effect: training a frontier model on Ascend clusters takes meaningfully more engineering effort than on equivalent Nvidia hardware. Chinese companies are absorbing this cost because they have no alternative at the high end. Some, like Alibaba's Qwen team, have published optimizations that suggest the overhead is narrowing.

This has an important implication for Western analysis. The market share shift is real and structural. The capability parity claim is not proven. Conflating them produces bad conclusions.

DeepSeek, Qwen, and What They Are Actually Running On

Two Chinese model families get the most attention in Western AI circles right now: DeepSeek, from the Shenzhen-based High-Flyer Capital Management research lab, and Qwen, from Alibaba's Hangzhou infrastructure teams.

DeepSeek is the more complicated case. The DeepSeek-V3 technical report is explicit: training used 2,048 H800 GPUs. That hardware was purchased before October 2023. DeepSeek's cost efficiency story — the widely cited $5.5 million training run figure — is real, but it runs on pre-export-control Nvidia hardware. Claims that DeepSeek's frontier models were trained on Ascend are unverified. If anyone tells you DeepSeek R2 or a successor model is fully Ascend-trained, ask for the source before accepting it. The question of what DeepSeek uses for inference is less publicly documented. For context on where DeepSeek is pushing its infrastructure, the work on DSec sandbox infrastructure for agent training is worth reading alongside the chip story.

Qwen is the cleaner Ascend case. Alibaba has published Ascend-specific optimization documentation for Qwen 2.5 model training and inference. For Alibaba — which sells Ascend compute through its cloud business — this is strategically logical. The public documentation suggests it is not a checkbox claim.

The distinction matters: the frontier Chinese models that Western professionals read about were mostly trained on Nvidia hardware China stockpiled. The transition to Ascend-native frontier training is ongoing, not complete.

What This Changes for Western Founders and Consultants

If you track Chinese AI companies as competitors: The hardware base is shifting. Within two to three years, the frontier models emerging from Chinese labs will increasingly be trained on Ascend clusters — not because Ascend beat Nvidia in benchmarks, but because Nvidia is no longer an option. Whether Ascend infrastructure can sustain the same training run efficiency as H100 clusters is an open question with real implications for the capability trajectory of Chinese AI development.

If you buy or advise on AI infrastructure: The argument that "Chinese AI is stuck because of chip restrictions" is no longer a tenable planning assumption. The restrictions are real. The workarounds — primarily Ascend, with domestic inference chips like Biren and Cambricon filling adjacent roles — are real too. The question worth modeling is how large the productivity differential between Ascend and H100 clusters actually is, and whether it is closing.

If you evaluate Chinese AI tools and models: Qwen 2.5, Ernie 4.0, and Kimi's latest releases are production software running at scale inside China, increasingly on Ascend hardware. The chip is not a reason to trust or distrust the output quality — but it is context worth having when someone claims Chinese AI is hardware-constrained into irrelevance.

If you work in enterprise software with Chinese clients or partners: Data residency rules increasingly funnel Chinese AI workloads toward domestic cloud providers running domestic chips. This is a procurement and compliance consideration, not just a geopolitical talking point.

When NOT to Conclude Huawei Has Won

Don't assume market share equals capability parity. Huawei's dominance in new AI accelerator shipments inside China is structural — a consequence of removing the competition, not of surpassing it. The performance gap between Ascend 910B and H100 is not publicly documented at the rigor required to make confident parity claims.

Don't assume China's AI development is now fully autonomous. SMIC, which manufactures some Ascend chips, faces its own equipment restrictions under export controls covering advanced semiconductor tooling. Yields at the most advanced process nodes China can access remain a real constraint. Huawei has found creative responses, but the manufacturing dependency has not been resolved.

Don't assume DeepSeek's results predict what Ascend-native training produces. The H800 clusters that trained DeepSeek-V3 represent hardware acquired under an earlier regulatory regime. The first frontier models trained entirely on Ascend infrastructure, at full scale, have not yet shipped publicly. That transition is happening and worth watching closely.

Don't extrapolate from one city. Beijing-based labs (Baidu, Zhipu) have different hardware relationships than Shenzhen-based research teams or Hangzhou-based Alibaba. China's AI ecosystem is not one thing.

Where This Is Heading

The Ascend 910C is in the pipeline. Huawei has signaled a next-generation chip beyond the 910B, but specs are unconfirmed and no shipping date has been officially announced. If the 910C ships at scale with material performance gains, the capability gap conversation changes, and Chinese analysts widely expect it to be the decisive product.

The software ecosystem will keep closing. CANN's maturity gap with CUDA is real, but Alibaba, ByteDance, and Baidu are collectively putting substantial engineering resources into it. "The software is not production-ready" is already an outdated claim. The question now is how close to CUDA parity it gets, and how fast.

Inference is a separate, faster-moving story. While training gets the attention, Chinese companies have been deploying inference chips from multiple domestic vendors at significant scale. Cambricon in particular has built a strong position in inference silicon. The training chip market and the inference chip market are not the same, and Huawei is not equally dominant in both.

The regulatory escalation is not finished. Each round of U.S. export tightening has structurally strengthened Huawei's domestic position. If controls extend further — into cloud services or chip manufacturing equipment — the dynamic shifts again. Treat this as a moving target for analysis, not a settled outcome.

The real test is in the next generation. Five years from now the meaningful question will not be whether Ascend caught H100. It will be whether the ecosystem built on top of Ascend — the compilers, the libraries, the model optimization toolchains — is deep enough to sustain frontier AI development through the next capability threshold. Alibaba's Qwen work offers early evidence; the answer remains cautiously open.

FAQ

Has Ascend 910B matched the H100's performance? Not to any independently verified standard. The H100 remains the benchmark for AI training chip performance globally. The 910B is a credible alternative for large-scale training on Chinese domestic clusters, and major Chinese model labs have shipped real products with it. But published performance comparisons come from first-party sources without independent replication. Close enough for production use; unproven at H100 parity.

Why can't Chinese companies just use cloud compute from U.S. providers? Some did historically. The export controls now cover not just chip sales but access to cloud compute using covered chips. Chinese companies building frontier AI are increasingly unable to purchase H100 compute time from U.S. cloud providers under the current regulatory framework. Domestic compute is increasingly a legal requirement, not just a preference.

Is Nvidia completely out of China's AI market? Not entirely. Lower-tier Nvidia chips — certain professional and workstation cards that fall below the export control performance thresholds — remain legal for sale. These are used for inference, developer workstations, and lighter training jobs. The high-end AI training cluster market, where large-scale model development happens, is effectively closed to Nvidia's best hardware.

What does this mean for the capability gap between Chinese and Western AI? That is the right question, and the honest answer is: unclear. The chip restriction was intended to slow Chinese AI development by limiting access to advanced hardware. Whether it has succeeded depends on how large the Ascend productivity gap is and how fast it is closing — and that data is not publicly available at the precision required to make a confident call either way.

If I'm evaluating a Chinese AI vendor's platform, should I care what chips they run on? For most commercial use cases — APIs, SaaS tools, B2B software — the underlying chip is not operationally relevant. Output quality, latency, and pricing are what matter. The chip question becomes relevant if you are doing technical due diligence on a Chinese AI infrastructure company, evaluating a cloud partnership, or assessing the long-term capability trajectory of a Chinese model lab you compete with or rely on.

What about Biren and Cambricon? They matter, but in a different tier. Biren (Shanghai) and Cambricon (Beijing) are building domestic AI chips targeting specific workloads, with Cambricon holding a particularly strong inference position. Neither is competing with Ascend in high-end training cluster deployments. They matter for the ecosystem story: China is building redundancy into its AI hardware supply chain, not concentrating the entire bet on Huawei.

Is this story about AI or about geopolitics? Both, inseparably. The export controls are a geopolitical instrument. The technology response is real engineering. Anyone who tells you this is purely about national security is not explaining why Chinese model labs are investing millions in Ascend optimization toolchains. Anyone who tells you it is purely about technology is not explaining why the entire market structure changed within weeks of an October 2023 regulatory update.

D
Debby Wang is BestAIFor's China AI Correspondent, covering the tools, startups, and policy shifts coming out of China's AI ecosystem. Based in Shenzhen, she writes for Western founders and professionals who want to understand what's actually happening - without the hype or the panic. Her focus areas include physical AI, robotics, medical applications, AI hardware, and the social and legal impact of automation.

Related Articles