The open-source language, built on Apache TVM, compiles GEMM and attention kernels toward NVIDIA, AMD, Apple and Huawei hardware. DeepSeek’s TileKernels library raised its requirement to this release the same day.
TileLang reached version 0.1.15 on Sept. 30, 2026, with a native backend for Huawei’s Ascend 950 neural processing unit. Maintainer Lei Wang published the release on the tile-ai GitHub repository. The notes describe an end-to-end path from Python kernel code to NPU binaries.
The project is not a new chip vendor. It is a language.
Tile Language, also called tile-lang, is “a concise domain-specific language designed to streamline the development of high-performance GPU/CPU/NPU kernels (e.g., GEMM, Dequant GEMM, FlashAttention, LinearAttention),” the repository states. It uses “a Pythonic syntax with an underlying compiler infrastructure on top of TVM” so developers can “focus on productivity without sacrificing the low-level optimizations necessary for state-of-the-art performance.”
Apache TVM is the compiler underneath. TileLang does not replace CUDA, ROCm or Huawei’s CANN toolkit. It sits above them. A kernel is ordinary Python until a decorator turns it into TVM intermediate representation. Arguments carry shape and data type. A Kernel block launches a grid of tile programs. The compiler then binds that tile program to whatever backend the machine exposes.
Researchers described the split in a paper submitted to arXiv on April 24, 2025, and revised on April 27. “TileLang: A Composable Tiled Programming Model for AI Systems” lists Lei Wang, Yu Cheng and Yining Shi as equal contributors, with further authors from Peking University, Imperial College London and Microsoft Research. The abstract says TileLang “decouples scheduling space (thread binding, layout, tensorize and pipeline) from dataflow, and encapsulated them as a set of customization annotations and primitives.” Users, the authors wrote, can “focus on the kernel’s data-flow itself, while leaving most other optimizations to compilers.”
The evaluation, they wrote, “shows that TileLang can achieve state-of-the-art performance in key kernels.” The abstract does not extend that claim to every operator or every accelerator.
Version 0.1.15 is the hardware increment. Release notes call the new work “an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming.” The package adds tilelang.ascend.language and a target named ascend for the Ascend 950, marked in the notes as dav-3510. One kernel can combine Cube matrix math and Vector computation, using SimdVF and SimtVF regions. Programmers can name Unified Buffer, L1 and L0 storage, issue tiled copies and cross-core transfers, and write MXFP8 and MXFP4 block-scaled GEMM. Examples cover GEMM, DeepGEMM-style kernels, FlashAttention forward and backward, RMSNorm and FP8 quantization. “Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects,” the notes say.
DeepSeek shipped a matching note the same day. Its TileKernels repository, which showed about 1,900 stars, calls itself “a library of dozens of highly optimized kernels implemented in TileLang, a domain-specific language supporting multiple hardware backends.” The kernels cover mixture-of-experts routing, Engram, quantization and manifold hyper-connections. “Most kernels achieve performance close to the hardware’s compute or memory bandwidth limits,” the repository states. “All of these kernels have already been used in our internal training and inference workloads.”
The Sept. 30 entry says Huawei Ascend support was added and that, “Following the NVIDIA path, the kernels now ship a second backend that is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs.” TileKernels requires TileLang 0.1.15 or higher. Its CUDA path asks for an NVIDIA SM90 or SM100 GPU and CUDA Toolkit 13.1 or higher. The Ascend path asks for an Ascend 950 and CANN 9.2.0 or higher.
The main repository showed about 7,900 stars. Stable installs are a single command, pip install tilelang. Nightly wheels sit on a separate index. Current documentation lists NVIDIA CUDA from SM70 through SM120, AMD GPUs through ROCm, Apple Silicon through Metal, the new Ascend 950 path and an experimental LLVM backend for CPUs. An autotuner can search block sizes and pipeline stages, compile candidates in parallel and cache the fastest legal result. None of those listings is a speed claim. They mark where the compiler has been aimed.
For a kernel team, the release draws a line. Data movement stays in Python. Schedule details — thread binding, layout, pipeline — can be annotated or left to the compiler. NVIDIA and AMD remain the mature GPU targets. Ascend 950 is the new NPU target, with storage still named by the programmer and synchronization inserted by the backend. DeepSeek’s library is the public sign that those kernels are already inside a training stack.
Lock-in math, for anyone buying accelerators, is simpler than the compiler paper. One language. One Sept. 30 tag. Two vendors in the DeepSeek install notes.

