Wednesday’s open-source update posts near-peak sparse attention on Huawei’s newest NPU, retunes a fused decoding kernel and breaks compatibility with earlier DeepSeek models.
DeepSeek on Wednesday released sparse attention kernels for Huawei’s Ascend 950 processor inside FlashMLA, the company’s open library of attention code used to run DeepSeek-V4.1.
The same drop removes Hopper support. It also retires kernels for earlier models and changes how FP8 and FP4 key-value caches are stored. Teams still running DeepSeek V3, V3.2 or V4.0 will need an older commit.
FlashMLA is not a model. The repository describes it as “DeepSeek’s library of optimized attention kernels, powering inference of the DeepSeek-V4.1 model on NVIDIA GPUs and Huawei Ascend NPUs.” The repo holds token-level sparse attention for prefill and for decoding with an FP8 or FP4 KV cache, a fused kernel and dense attention kernels for prefill and the backward pass.
That split matters for anyone buying inference capacity. Attention, not the parameter count alone, is what keeps reading memory once a long prompt is already sitting in cache.
The Sept. 30 note is direct. “We’ve released sparse attention prefill and decoding kernels for the Huawei Ascend 950 NPU, which achieve up to 410 TFlops (95% hardware peak) and 360 TFlops (83% of hardware peak) during prefill and decoding, respectively,” the project documentation states. A technical report on the algorithms and the optimization work shipped with those kernels. DeepSeek also reported that it improved the fused norm-RoPE-attention-RoPE-cast kernel under decoding settings by around 10% to 15%.
NVIDIA figures sit in the same documentation, measured on a B200 with CUDA 13.3. The fused kernel reaches up to 1,460 TFlops in prefill and 950 TFlops in decoding. MLA prefill hits up to 1,350 TFlops. MLA decoding hits up to 1,024 TFlops. Dense multi-head attention prefill on that chip is listed at up to 1,460 TFlops forward and 1,000 TFlops backward. An April 22, 2025, update had put compute-bound MLA decoding at up to 660 TFlops on an H800 SXM5, a 5% to 15% gain over the prior kernel, with an interface the team said was fully compatible at the time.
The break with older code is explicit. “In the 2026.09.30 release, we removed support for the Hopper architecture and for earlier models (including DeepSeek V3 / V3.2 / V4.0), and we changed the FP8 / FP4 KV cache format. This version is therefore not compatible with previous ones,” the README states. Operators who still need those models, or the old cache layout, are pointed to a prior commit.
What the library actually runs is narrower than the name suggests. Sparse kernels power DeepSeek Sparse Attention, the mechanism introduced with DeepSeek-V3.2. Those kernels arrived Sept. 29, 2025, with published peaks of 640 TFlops in prefill and 410 TFlops in decoding, plus a write-up on the FP8 sparse decoding kernel. Kernels aimed at DeepSeek-V4.1, covering prefill and decoding with FP8 or FP4 cache, followed on Sept. 10, 2026, alongside the fused operator. Q-norm inside that fused kernel is used in V4, not in V4.1. The fuse covers Q-norm, Q-RoPE, core attention, conjugate O-RoPE and the cast to FP8, cutting the overhead of launching those small kernels separately. Decoding support stops at FP8 and FP4. Unquantized bfloat16 KV cache is not supported.
Cache layout is now part of the contract. For V4.1, format is read from the last dimension of the key cache: 528 bytes per token for FP8, 288 bytes per token for FP4. In the FP8 layout, the first 512 bytes hold float8_e4m3 values across all 512 dimensions, RoPE dimensions included, and the last 16 bytes hold float8_e8m0 scales, one for every 32 values. The FP4 layout packs 512 e2m1 values into 256 bytes, two values per byte, and stores float8_e4m3 scales in the remaining 32 bytes, one for every 16 values. No bfloat16 slice remains.
Hardware gates are tight. NVIDIA builds target SM100 and SM103 GPUs, CUDA 13.1 or newer, with 13.2 or newer recommended, and PyTorch 2.0 or newer. Ascend builds need an Ascend 950 NPU, CANN 9.2.0 or newer, torch_npu and PyTorch 2.0 or newer. Hopper is out of this release. Install is a clone, a submodule update and a pip install. The build target can be forced to CUDA or Ascend.
Jiashi Li, Shengyu Liu and Yuanhang Sun are listed as authors on the project citation. The kernel work sits downstream of FlashAttention-style online softmax scheduling. The April 2025 deep-dive described a “seesaw” schedule meant to overlap CUDA cores, tensor cores and memory movement on a decoding kernel that, at DeepSeek’s head count, is compute-bound rather than memory-bound. An Aug. 1, 2025, note credited an NVIDIA pull request for dense multi-head attention forward and backward kernels on SM100.
For an inference lead, the practical read is simpler than the kernel names. Near-peak sparse attention on Ascend 950 is now public code. Blackwell-class NVIDIA numbers are published beside it. The cost of taking the update is a clean break with Hopper and with every DeepSeek model before V4.1.

