The MIT-licensed library, used in DeepSeek V3.2, V4 and V4.1, claims a 2 to 20 times gain over PyTorch’s stock topk. CUDA support arrived Sept. 10. Huawei Ascend kernels followed Sept. 30.
DeepSelect is not a model. It is a TopK kernel.
DeepSeek published it for one job that shows up twice in inference: pick the relevant positions inside DeepSeek Sparse Attention, and pick candidate tokens when the sampler faces a large vocabulary. The project page is blunt. “DeepSelect is a high performance implementation of the TopK kernel used in DeepSeek Sparse Attention (DSA) (which is used in DeepSeek V3.2, DeepSeek V4, and DeepSeek V4.1 models) and the sampler.”
Platforms are NVIDIA CUDA and Huawei Ascend. The same paragraph says the kernel “achieves 2 ~ 20x speedup compared to vanilla torch.topk.” That figure is a range, not a promise tied to one chip or one batch size. The repo’s own benchmark script reports the ratio against torch.topk on the same input, and it scores effective memory bandwidth. “TopK does no floating-point math, so a FLOP rate would not be meaningful here.”
Version 1.0.0 and a short algorithm note went up Sept. 10. On Sept. 30 the news line read, “We’ve released TopK kernels for Huawei Ascend NPU.” Same day, DeepSeek said on its WeChat account that it was open-sourcing programming infrastructure for Huawei’s Ascend platform, including compute and communication libraries, Reuters reported. Accounts of that drop place DeepSelect with TileLang, DeepGEMM, DeepEP, TileKernels and FlashMLA. DeepSelect is the selection piece. The others cover language, matrix math, communication and attention.
Why rewrite TopK at all? The lab’s deep-dive answers it in three sentences. “In DeepSeek Sparse Attention (DSA), TopK is used to select the most relevant positions from a large number of context tokens.” Then: “During sampling, TopK is also used to select candidate tokens from a large vocabulary.” And the motive: “To improve end-to-end inference performance, we designed a new TopK algorithm.”
The method keeps a threshold, T, starting at negative infinity. Input is read once, in contiguous blocks of size B, and those blocks are walked in random order so a nasty layout does not stall the run. Values above T land in a candidate buffer. When that buffer reaches k plus a second block size, B2, a radix select cuts it back to k and T moves up. It does not fall. For k of 512, the note calls B = B2 = 1024 a reasonable setting. When both block sizes sit on the order of k, the extra work is O(k log(N/k)), far smaller than the input length N. Extra space is O(k + B + B2), meant to live in shared memory. The CUDA path is CUDA C++, with PTX and bit operations aimed at the filter step.
Coverage is narrow on purpose. The Lightning Indexer case, the sparse-attention path, wants torch.bfloat16 and is tuned mainly for topk of 512, which the README ties to DeepSeek V4. Topk must be 4,096 or less. Batch size and vocabulary length are treated as open. Sampling wants torch.float32, a vocabulary around 128,000 — the benchmark uses 129,280 — and the same 4,096 cap. One split is hard. “Currently, only the CUDA implementation supports torch.float32 as the input dtype. The Ascend implementation supports torch.bfloat16 only.” Ascend indices are int32 only. CUDA can emit int32 or int64.
Install is a clone, a submodule update and pip install -v . Rows must be contiguous on the last dimension, and the input stride has to match deep_select.get_stride_requirement. Unaligned inputs need padding. Two switches change the cost. “Disable sorted_index unless the output has to be ordered by index; enabling it costs performance.” And: “Set return_value=False when the values are not needed. This skips the value output and is faster.” The usage note marks that skip at about 10 percent. Sorting by value is limited to the float32 CUDA path.
NaNs are not waved through. “NaN checking is always on. With the default abort_when_nan_found=True the kernel invokes trap() and aborts.” On CUDA, rows no longer than topk skip that check.
The citation names Yi Qian, Shengyu Liu and Yichen Li. No paper sits next to the code. The repository is under the MIT License. The deep-dive is the documentation.
For an inference lead, the question is plain. Is selection eating time on V3.2, V4 or V4.1, and is the box NVIDIA or Ascend? DeepSelect answers that slice. It does not replace a serving stack.

