The MIT-licensed port keeps the DeepGEMM programming interface and, in company tests on Ascend 950DT hardware, reports dense matrix-multiply utilization as high as 99.8%.
DeepSeek on Wednesday published DeepGEMM-Ascend, an open-source matrix-multiplication kernel library written for Huawei Technologies Ascend neural processing units and documented for the Ascend 950 series.
The Hangzhou company posted the code the same day it said, on its official WeChat account, that it was open-sourcing programming infrastructure for Huawei’s Ascend platform. The GitHub news log dates the first DeepGEMM-Ascend release to Sept. 30, 2026. The kernels “are designed to achieve near-peak hardware performance,” the repository said.
DeepGEMM-Ascend is a port of DeepGEMM, DeepSeek’s existing kernel library for Nvidia GPUs. The README states that the Ascend package is fully API-compatible with DeepGEMM and supports BF16, FP8 and FP4 general matrix multiplication, or GEMM, along with multi-query attention logits and MegaMoE. Users install one package and keep the same interfaces and workflow used on other supported platforms, the project said.
That compatibility stops short of identical data layouts. A project note says the scaling-factor format on Ascend differs from Nvidia’s: each pair of UE8M0 scaling factors along the K dimension is packed into an int16, and the packed values are stored in MN-major order. Teams moving quantized workloads still have to adapt that layout.
The library wraps Ascend matrix multiply-add primitives and is meant to hide fractal layouts, alignment rules, address math and bulky low-level parameters. It uses Ascend-specific methods such as sparse data loading and coroutine-based pipelining. Project leads listed in the repository are Kexing Zhou, Zhean Xu and Chenggang Zhao.
Company benchmarks were run on an Ascend 950DT with CANN 9.20, using bench_msprof and a cold L2 cache. Shapes follow the DeepGEMM test suite and cover DeepSeek model-series training and inference cases, the README said. Those figures have not been independently verified.
On a 4096-by-7168-by-16384 dense GEMM, DeepSeek reported 1,701 teraflops for FP4-by-FP4, or 98.3% of a listed 1,730-teraflops hardware limit; 861 teraflops for both FP8-by-FP4 and FP8-by-FP8, or 99.5% of 865 teraflops; and 431 teraflops for BF16-by-BF16, or 99.8% of 432 teraflops. Latency on the BF16 case was listed at 2,229.9 microseconds.
MQA scoring for DeepSeek’s Lightning Indexer is described as FIX-pipe bound, saturating that pipe at 99% utilization. MegaMoE fuses expert-parallel dispatch, two grouped GEMMs, SwiGLU and combine. Over EP8 with top-k of 6 and one shared expert, a 384-expert case with hidden size 7,168, intermediate size 3,072 and 16,384 tokens was listed at 846.3 teraflops.
Build requirements include an Ascend NPU validated on the 950 series, CANN 9.20 with the bisheng compiler and ld.lld linker, the torch_npu package, Python 3.10 or higher, and a C++20 toolchain with format-library support. Developers clone the repository with submodules, run develop.sh, then install with pip and no build isolation. The code is released under the MIT License.
DeepGEMM-Ascend sits inside a broader Wednesday drop. DeepSeek also pointed developers to TileLang, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect. In the WeChat post, the company said Huawei gave full support and that the two firms jointly advanced a supernode built on 128 Ascend 950 chips while tuning compute and communication.
“To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware’s full performance potential,” DeepSeek said. “TileLang was created precisely to meet this need.”
The company said TileLang offers “a simpler programming model” than Nvidia’s CUDA. Bloomberg, citing the same WeChat post, quoted DeepSeek as saying Huawei provided “unreserved and vigorous support.”
For executives, the release is a software path onto Ascend 950 iron that does not require a new Python API for teams already on DeepGEMM. For kernel engineers, the tests, shapes and source are in the tree. What remains unshown is how those kernels behave on other Ascend generations, under production schedulers, or in third-party hands.
No tracking, no middleman. Follow by RSS (nothing is collected) — or add your email to our self-hosted list.
RSS feed →
