DeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
DeepSeek released DeepEP-Ascend and DeepGEMM-Ascend, open-source communication and GEMM libraries for Huawei Ascend NPUs. DeepEP-Ascend achieves 373–375 GB/s dispatch and 345–347 GB/s combine at EP8 on Ascend 950DT. DeepGEMM-Ascend reaches up to 99.8% of hardware limit for BF16 GEMM.
China context
- Original name
- 深度求索
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can use these libraries to optimize MoE training and inference on Huawei Ascend NPUs, with APIs aligned to NVIDIA DeepEP and DeepGEMM.
- For investors
- The release indicates DeepSeek's commitment to Huawei's Ascend ecosystem, increasing the viability of Ascend-based AI infrastructure for large-scale deployments.
Translated from Chinese. Quotes and facts link to the original sources.
DeepSeek published two new open-source repositories on GitHub: DeepEP-Ascend, a high-performance communication library for MoE training and inference on Huawei Ascend NPUs, and DeepGEMM-Ascend, a matrix multiplication kernel library ported to Ascend. DeepEP-Ascend provides expert-parallel all-to-all operations with FP8 dispatch and deferred epilogues, plus experimental pipeline parallelism, context/data parallelism, and remote memory access primitives. DeepGEMM-Ascend supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE, with API compatibility to the original DeepGEMM. Performance measurements on Ascend 950DT with CANN 9.2.0 show DeepEP-Ascend dispatch bandwidth of 373–375 GB/s at EP8 and DeepGEMM-Ascend achieving up to 99.8% of hardware limit for BF16 GEMM. Both libraries are available on GitHub under the deepseek-ai organization.
DeepEP-Ascend uses HCCL/HCOMM, UBMEM, and URMA for communication, with kernels compiled at runtime via DeepJIT. DeepGEMM-Ascend abstracts Ascend MAD primitives and employs sparse data loading and coroutine-based pipelining to approach hardware limits. The libraries target Ascend 950 series NPUs and require CANN 9.2.0, Python 3.10+, and PyTorch/torch_npu.
DeepSeek's release of Ascend-optimized libraries signals growing support for Huawei's AI hardware ecosystem among major Chinese AI labs. This may accelerate adoption of Ascend NPUs for large-scale model training and inference, reducing reliance on NVIDIA GPUs in China.
For developers and enterprises using Huawei Ascend hardware, these libraries provide production-ready communication and GEMM kernels that can improve training and inference efficiency for MoE models. The open-source nature allows customization and integration into existing pipelines.
DeepEP-Ascend's README notes that Huawei's Q3 commercial HDK for Atlas 850E is planned for public availability around October 15, 2026, which may improve performance further. Ongoing work includes Bucket collectives, expert load balancing, and PP/Engram interfaces.