DeepSeek open-sources operator toolchain with native Huawei Ascend support, Leiphone reports
Reported by Leiphone · not yet confirmed by the company or a second independent outlet. We update this page when it is.
DeepSeek released open-source versions of its operator development tool TileLang and five core compute/communication libraries (DeepGEMM, DeepEP, FlashMLA, TileKernels, DeepSelect) with native support for Huawei Ascend 950 NPU, Leiphone reported on 2026-09-30.
China context
- Original name
- 深度求索
- Outside China
- Open weights · leiphone.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can evaluate the open-source TileLang and associated libraries for cross-platform operator development, potentially reducing dependence on CUDA for targeting Huawei Ascend hardware.
- For investors
- The release may accelerate adoption of Huawei Ascend NPUs in AI infrastructure, which could shift demand away from Nvidia in markets where Chinese chips are available.
Translated from Chinese. Quotes and facts link to the original sources.
DeepSeek has ported its full operator development toolchain, originally built for Nvidia GPUs, to Huawei's Ascend 950 NPU. The release includes TileLang, a core operator development product, now natively supporting Ascend, along with DeepGEMM, DeepEP, FlashMLA, TileKernels, and DeepSelect. This marks the first time Chinese-made compute has achieved capability parity with Nvidia's ecosystem at the core operator tool level, according to Leiphone.
The release targets the high barrier of writing AI operators, where manual optimization is needed to raise hardware utilization from around 20% to over 95%. TileLang and the associated libraries aim to provide a cross-platform toolchain that reduces the need for hand-written CUDA code and enables code portability across different chip architectures.
Developers using Nvidia's CUDA ecosystem now have an open-source alternative that runs on Huawei Ascend NPUs, potentially lowering the cost and complexity of porting AI workloads to Chinese-made chips. This directly affects teams seeking to avoid CUDA lock-in or to deploy on domestic Chinese hardware.
For enterprises and developers, this toolchain could reduce the engineering effort required to optimize AI models for Huawei Ascend NPUs, potentially making Chinese-made compute more accessible for AI training and inference. However, the practical benefits depend on real-world adoption and performance verification.
A key signal to watch is whether independent developers adopt TileLang for production workloads on Ascend hardware and whether performance benchmarks match or exceed those of Nvidia's native toolchain. The absence of third-party validation means the claimed parity remains company-reported.