The TileLang kernel library for mixture-of-experts routing, quantization and Engram now runs the same Python APIs on Nvidia GPUs and Huawei NPUs.
Hangzhou-based DeepSeek on Wednesday published Huawei Ascend support in TileKernels, an open-source library of training and inference kernels written in TileLang. The repository’s news note is dated Sept. 30, 2026. It says the kernels now ship a second backend that is selected automatically at runtime, so the same Python interfaces run on Nvidia GPUs and Huawei NPUs.
TileKernels is not a demo shelf. The README describes it as a library of dozens of highly optimized kernels for large-language-model work: mixture-of-experts routing, Engram, quantization and manifold hyper-connections. “Most kernels achieve performance close to the hardware’s compute or memory bandwidth limits,” the project says. “All of these kernels have already been used in our internal training and inference workloads.”
The code is implemented in TileLang, a domain-specific language with Python-like syntax and a compiler stack built on Apache TVM. TileLang’s maintainers list Nvidia CUDA as the primary backend and also document support for AMD ROCm, Huawei Ascend 950, Apple Metal and several experimental targets.
Wednesday’s README addition is short and operational. “Added Huawei Ascend support and updated the usage documentation,” it says. “Following the NVIDIA path, the kernels now ship a second backend that is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs.” Developers do not maintain two call sites for the listed operators. The runtime picks the backend.
That change sits inside a broader DeepSeek release the company posted on its official WeChat account the same day. Reuters reported that DeepSeek said it was open-sourcing programming infrastructure for Huawei’s Ascend platform, including compute and communication libraries, and that Huawei provided support for the work. The two companies also advanced a “supernode” configuration based on 128 Ascend 950 chips.
DeepSeek’s case for TileLang, as quoted by Reuters, is about writing kernels once and still pressing the metal. “To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware’s full performance potential.” The company added: “TileLang was created precisely to meet this need.” Reuters also reported that DeepSeek said TileLang offers “a simpler programming model” than Nvidia’s CUDA.
Chinese-language reports on the WeChat post listed companion repositories released or updated for Ascend: DeepGEMM for matrix math, DeepEP for cross-device communication, FlashMLA for sparse attention, DeepSelect for data selection, and TileKernels for the common vector and memory-access operators used in training and inference.
Inside TileKernels, the feature list is specific. MoE routing covers top-k expert selection and scoring. Quantization covers per-token, per-block and per-channel FP8 and FP4 cast and dequantization, plus fused SwiGLU and quantization ops. Engram kernels fuse RMSNorm, expose forward and backward passes, and reduce weight gradients. Manifold HyperConnection kernels include Sinkhorn normalization and mix splitting and application. There is a rotary position embedding kernel under transform, a random-number kernel, and a torch.autograd.Function wrapper for the Engram gate.
The tree matches that split: moe/, quant/, engram/, mhc/, transform/, rand/, modeling/, plus PyTorch reference code and test helpers.
Hardware cuts off older cards. The documented CUDA path needs an Nvidia SM90 or SM100 GPU and CUDA Toolkit 13.1 or higher. The Ascend path needs an Ascend 950 NPU and CANN 9.2.0 or higher. Software floors in the current README are Python 3.12, PyTorch 2.13 and TileLang 0.1.15. The published package metadata still lists looser pins and marks the project Alpha.
Install from PyPI with pip install tile-kernels, or from a clone with pip install -e ".[dev]". Tests run through pytest. A single file can be checked for correctness or with --run-benchmark. TK_TEST_LEVEL switches core tests and a fuller suite.
The repository is released under the MIT License. A README acknowledgement thanks TileLang’s developers and Huawei “for its technical support and engineering expertise throughout the development of Tile Kernels’ Ascend backend.” The suggested citation lists Xiangwen Wang, Chenhao Xu, Huanqi Cao and co-authors, dated 2026.
For engineering leads, the practical claim is narrow. One library. Two backends. Operators already used in DeepSeek’s own training and inference, now callable on Hopper- and Blackwell-class Nvidia silicon or on Ascend 950 through the same Python surface.

