Event date · · DeepSeek

DeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs

DeepSeek 深度求索Chinese AIOpen weights
FACT STATEMENT

DeepSeek released DeepEP-Ascend and DeepGEMM-Ascend, open-source communication and GEMM libraries for Huawei Ascend NPUs. DeepEP-Ascend achieves 373–375 GB/s dispatch and 345–347 GB/s combine at EP8 on Ascend 950DT. DeepGEMM-Ascend reaches up to 99.8% of hardware limit for BF16 GEMM.

China context

Original name
深度求索
Outside China
Open weights · github.com
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can use these libraries to optimize MoE training and inference on Huawei Ascend NPUs, with APIs aligned to NVIDIA DeepEP and DeepGEMM.
For investors
The release indicates DeepSeek's commitment to Huawei's Ascend ecosystem, increasing the viability of Ascend-based AI infrastructure for large-scale deployments.

Translated from Chinese. Quotes and facts link to the original sources.

What happened

DeepSeek published two new open-source repositories on GitHub: DeepEP-Ascend, a high-performance communication library for MoE training and inference on Huawei Ascend NPUs, and DeepGEMM-Ascend, a matrix multiplication kernel library ported to Ascend. DeepEP-Ascend provides expert-parallel all-to-all operations with FP8 dispatch and deferred epilogues, plus experimental pipeline parallelism, context/data parallelism, and remote memory access primitives. DeepGEMM-Ascend supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE, with API compatibility to the original DeepGEMM. Performance measurements on Ascend 950DT with CANN 9.2.0 show DeepEP-Ascend dispatch bandwidth of 373–375 GB/s at EP8 and DeepGEMM-Ascend achieving up to 99.8% of hardware limit for BF16 GEMM. Both libraries are available on GitHub under the deepseek-ai organization.

Technical significance

DeepEP-Ascend uses HCCL/HCOMM, UBMEM, and URMA for communication, with kernels compiled at runtime via DeepJIT. DeepGEMM-Ascend abstracts Ascend MAD primitives and employs sparse data loading and coroutine-based pipelining to approach hardware limits. The libraries target Ascend 950 series NPUs and require CANN 9.2.0, Python 3.10+, and PyTorch/torch_npu.

Industry impact

DeepSeek's release of Ascend-optimized libraries signals growing support for Huawei's AI hardware ecosystem among major Chinese AI labs. This may accelerate adoption of Ascend NPUs for large-scale model training and inference, reducing reliance on NVIDIA GPUs in China.

Decision value

For developers and enterprises using Huawei Ascend hardware, these libraries provide production-ready communication and GEMM kernels that can improve training and inference efficiency for MoE models. The open-source nature allows customization and integration into existing pipelines.

What to watch

DeepEP-Ascend's README notes that Huawei's Q3 commercial HDK for Atlas 850E is planned for public availability around October 15, 2026, which may improve performance further. Ongoing work includes Bucket collectives, expert load balancing, and PP/Engram interfaces.

Latest in Chinese AI

  1. AlibabaAlibaba DAMO Academy Unveils DAMO EAGLE AI Model for Early Esophageal Cancer Detection
  2. Manus AIManus AI releases Manus 2.0
  3. AlibabaAlibaba Unveils Agentic Computer, AI Wearables and More at 2026 Apsara Conference
  4. AlibabaAlibaba Launches Qwen Intelligence to Power Next-Generation Agentic Smartphones
  5. AlibabaAlibaba Releases Qwen3.8-Omni-Flash with 1M-Token Context for Audio-Visual Agents

All China AI Events

CHINA AI WEEKLY

Get the week in Chinese AI, in English.

One weekly issue of verified model, company, robotics and policy changes, each with its original source and outside-China availability.