DeepSeek published six open-source software modules for Huawei’s Ascend accelerators on 30 September, announcing them through its WeChat account. The package includes a version of TileLang, the company’s high-level chip programming language, adapted to run on the Ascend 950 — and DeepSeek is explicit that it is positioning it against Nvidia’s CUDA.
The other five are the pieces a training stack actually needs: DeepGEMM for matrix multiplication, FlashMLA for attention, TileKernel, DeepSelect and DeepEP. Reuters reported that the two companies also worked together on a “supernode” built from 128 Ascend 950 chips, optimised for both computation and the communication between them.
Software is the moat, not the die
Nvidia’s durable advantage has never been that its chips are the only fast ones. It is that CUDA is nineteen years old, that every framework targets it, and that the accumulated kernels, profilers and workarounds are worth more than any single generation of hardware. Chinese accelerators have existed for years; the reason they stayed in the lab is that porting a serious training run to them cost more engineer-months than the chips saved.

DeepSeek’s pitch is aimed exactly there. It describes TileLang as offering “a simpler programming model” than the alternative, and frames the wider ambition as establishing “a high-level language that is universal, easy to program, and still capable of reaching the hardware’s full performance potential” — the foundational step, in its words, toward “a new generation of independent, self-controlled GPU software ecosystems”.
That is a political sentence as much as a technical one, and it should be read as both. It is also a claim about ergonomics that no independent party has yet tested.

Why DeepSeek is the company doing it
DeepSeek trains at frontier scale under export controls, which means it has already paid the cost of making constrained hardware work and has kernels written for it. Open-sourcing them converts a private workaround into shared infrastructure, and lowers the switching cost for every other Chinese lab that was waiting for someone else to go first. Huawei gets what it has needed since the Ascend line began: software written by people who train large models rather than by people who sell chips.
Nothing here has been benchmarked independently, and no throughput figures accompanied the release.
What to watch
Whether a lab outside DeepSeek publishes a training run on Ascend using these libraries, and whether the supernode’s 128-chip configuration shows up in any Chinese cloud’s public pricing. Adoption by someone with no reason to flatter either company is the only measurement that will settle this.