The open-source expert-parallel library aligns its public APIs with Nvidia DeepEP and reports 373–375 GB/s FP8 dispatch on eight Ascend 950DT ranks.
DeepSeek on Wednesday published DeepEP-Ascend, an open-source communication library for mixture-of-experts training and inference on Huawei Technologies Ascend neural processing units. The code lives on GitHub under the deepseek-ai organization.
The library handles the expert-parallel all-to-all path that mixture-of-experts models lean on. Tokens have to be dispatched to the experts that will process them, then combined after those experts finish. DeepEP-Ascend covers that dispatch and combine work, including FP8 dispatch and deferred epilogues. Public buffer APIs are aligned with the Nvidia edition of DeepEP. The parent project’s news note says the Ascend build offers the “Same API and full performance on HUAWEI Ascend 950 NPUs.”
Kernels are written in Ascend C. They ride Huawei’s HCCL and HCOMM stacks, plus UBMEM and URMA, and they compile at runtime through DeepJIT. Training, prefill and decode share one EPBuffer interface and the deep_ep Python package name. That matters to teams already shipping DeepEP on Nvidia silicon. They can keep the same buffer model instead of standing up a second communication API.
The drop sits inside a larger Wednesday release. DeepSeek said on its official WeChat account that it was open-sourcing Ascend infrastructure that also includes TileLang and related compute libraries. Reuters reported the Hangzhou firm’s language on that work.
“To build a new generation of independent, self-controlled GPU software ecosystem, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware’s full performance potential,” DeepSeek told Reuters. “TileLang was created precisely to meet this need,” the company said, adding that it offers “a simpler programming model” than Nvidia’s CUDA.
DeepSeek said Huawei gave “unreserved and vigorous support” during the research. The two companies jointly advanced a supernode design built on 128 Ascend 950 chips and jointly tuned compute and communication, according to the WeChat post and subsequent coverage.
DeepEP-Ascend’s own numbers are narrower and more cautious than that headline stack. Tests ran on Ascend 950DT NPUs with CANN 9.2.0 and a proof-of-concept hardware development kit that DeepSeek configured by hand. Traffic used netlayer 1 on the supernode’s external Clos network. Each rank held 16,384 tokens. Hidden size was 7,168. Routing was top-6 across 256 experts, with 32 AI cores and 64 AIVs, expert alignment of 128 and no bias. Dispatch used FP8 with row-major scales. Combine used BF16. Ten warmups. Fifty samples. Caches flushed.
Under those conditions, dispatch across all ranks measured 373–375 GB/s at EP8, 348–352 GB/s at EP16, 335–340 GB/s at EP32, 323–327 GB/s at EP64 and 313–320 GB/s at EP128. Combine ran 345–347, 338–341, 320–324, 294–298 and 272–278 GB/s on the same EP sizes. Timings include issue and drain. They exclude final epilogues.
Sustained dispatch reaches roughly 90–95% of the physical payload bandwidth limit for EP sizes up to 32, the README says. Larger EP sizes and combine are still under optimization. Combine carries extra local reduction cost and HBM contention with URMA. Huawei firmware changes are planned to ease that contention.
Those figures are not results from Huawei’s unreleased commercial kit. DeepSeek says the recommended public baseline is Huawei’s third-quarter commercial HDK for the Atlas 850E. Public availability is planned for mid-October 2026, around Oct. 15, through Huawei’s Atlas 850E software page. The date is a vendor plan. Users on earlier proof-of-concept kits may see lower bandwidth.
Hardware assumptions are strict. Linux on an Ascend host. Ascend 950 (A5) with UBMEM and UBC_CTP/URMA channels between ranks. An even number of ranks for multi-rank jobs. CANN with Ascend C, the Bisheng compiler, HCCL and HCOMM. Python 3.10 or newer, matching PyTorch and torch_npu, plus NumPy. A C++20 compiler that supports std::format. The validated stack is Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu and torch_npu 2.13.0rc1. Support on other Ascend generations or CANN versions has not been established by these measurements.
PP, Engram and Bucket collectives are experimental. Batched all-gather is in. Reduce-scatter and all-reduce kernels are still being built. Expert-load-balancing APIs are exposed; their Ascend communication kernels are not. Hybrid communication, CPU-backed Engram storage and graph capture are unsupported.
Install is ordinary enough once CANN is live: clone the repo, initialize the DeepJIT submodule, then python -m pip install --no-build-isolation .. Cited authors are Chenggang Zhao, Shangyan Zhou, Kexing Zhou, Rui Tian, Chenqi Zhao, Chenhao Xu, Yizhi Wang and Kuai Yu.
Companion Ascend repos published with the same stack include DeepGEMM-Ascend, TileKernels, FlashMLA and DeepSelect, plus TileLang.

