DeepSeek and Huawei's DeepGEMM Ascend: A Real CUDA Challenger?
DeepSeek and Huawei open-sourced DeepGEMM Ascend and Tile Lang, tools aimed at Nvidia's CUDA moat. Here's what they actually do.

What are DeepGEMM Ascend and Tile Lang?
DeepGEMM Ascend is an open-source library, built by DeepSeek, that runs fast matrix multiplication on Huawei’s Ascend AI chips. Tile Lang is a separate open-source programming language for writing AI kernels that compiles down to multiple chip backends, including Nvidia, AMD, Apple, and now Huawei Ascend. Together they’re part of a push to build a software layer for Chinese AI chips that doesn’t depend on Nvidia’s CUDA.
TL;DR
- DeepGEMM Ascend is DeepSeek’s port of its existing GEMM (general matrix multiplication) library to Huawei’s Ascend chips, keeping the same API so developers don’t need to relearn anything to switch hardware.
- Tile Lang is a Python-style kernel language, built on the TVM compiler framework, that already targets CUDA, AMD, and Apple Metal and now has a native backend for Huawei’s Ascend 950 chip.
- DeepSeek’s own benchmarks claim DeepGEMM Ascend hits close to 100% of the chip’s theoretical hardware limit on GEMM and mixture-of-expert workloads across BF16, FP8, and FP4 formats.
- The real target isn’t just Nvidia’s hardware, it’s CUDA’s software monopoly, estimated at around 4 million developers worldwide who are locked in by habit and tooling rather than by chip performance alone.
- Huawei provided technical backing during development and paired the software release with a super node system built from 128 Ascend 950 chips, suggesting a coordinated hardware-and-software push.
- High efficiency numbers show the software is well optimized, but they say nothing about how Ascend chips compare to Nvidia’s newest silicon in raw performance, that’s a separate and still open question.
- The release was quiet: no keynote, no big announcement, just a GitHub repo under an MIT license, with the code left for anyone to run and verify.
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
Why does Nvidia dominate AI hardware in the first place?
The usual answer is “better chips,” but that’s incomplete. Nvidia’s advantage is really a software advantage. For more than two decades, CUDA has been the default layer that tells Nvidia GPUs what to do. Developers learned it, built tools around it, and wrote millions of lines of code against it. The New York Times recently put the number of CUDA developers worldwide at an estimated 4 million.
That number matters more than any benchmark. A competitor can eventually build a chip that matches or beats Nvidia on raw specs. It’s far harder to replace 4 million people’s accumulated habits, debugging tools, documentation, and existing codebases. That’s the real moat, and it’s why Chinese chipmakers have struggled even as they’ve closed the hardware gap.
How does DeepGEMM Ascend actually work?
GEMM, general matrix multiplication, is the core operation inside every neural network: multiplying huge grids of numbers over and over, billions of times per training run or inference call. If that operation is slow, everything built on top of it is slow too.
DeepSeek already had a GEMM library for Nvidia GPUs called DeepGEMM. The new release ports that same library to Huawei’s Ascend chips while preserving the same API and package name. In practice, that means a developer already using DeepGEMM on Nvidia hardware can install the Ascend version and keep writing code the same way. The friction of switching chip families, normally one of the biggest reasons people stay on CUDA, is designed to disappear.
Underneath, the library wraps Ascend’s native matrix-multiply instructions in a thin layer that hides the usual low-level pain: memory layout rules, alignment, address calculations. That keeps the actual kernel code short and readable instead of requiring hardware-specific expertise. The library supports BF16, FP8, and FP4 number formats, along with attention-scoring operations and a fused mixture-of-expert operation DeepSeek calls Mega MoE, which is relevant for models like DeepSeek’s own and other mixture-of-expert architectures in the open ecosystem.
What do the benchmark numbers actually show?
According to DeepSeek’s own GitHub repository, the library gets very close to the Ascend chip’s theoretical hardware ceiling across several configurations: roughly 431 teraflops against a stated hardware limit of about 432 for one BF16 test, around 861 against a limit of 865 for FP8, and about 701 against 730 for FP4. The Mega MoE operation reportedly passed 800 teraflops on larger systems.
When a kernel runs that close to its hardware ceiling, it means the software isn’t wasting the chip’s capability. Almost all of what the silicon can physically do is being used. That’s a legitimate software engineering achievement.
But it answers only one question: is the software efficient? It does not answer a second, separate question: how does that hardware ceiling compare to Nvidia’s latest GPUs? Those are different measurements, and the DeepGEMM Ascend release only speaks to the first one. Without a direct, independent comparison against current Nvidia chips, the efficiency numbers say more about software quality than about whether Ascend chips are competitive with Nvidia’s best hardware.
Why does Tile Lang matter more than DeepGEMM Ascend in the long run?
DeepGEMM Ascend proves a specific chip can run specific operations fast. Tile Lang is an attempt at something bigger: a general-purpose way for ordinary developers to write fast AI kernels without becoming experts in any one chip’s low-level instruction set. That’s closer to what CUDA actually became over twenty years, not just a fast library, but a universal language layer that works across many kinds of code and lets a large developer population move fast.
Tile Lang lets developers write kernels in Python-style code describing how data moves in tiles and how matrix math executes, then compiles that down to hardware-specific fast code. It’s built on TVM, an existing open-source compiler framework, and already supported Nvidia CUDA, AMD, and Apple Metal before this release added a native backend for Huawei’s Ascend 950, including automatic scheduling and synchronization so developers don’t have to hand-tune every detail themselves. Notably, one of the kernels inside DeepGEMM Ascend (the one for DeepSeek’s hyperconnection module) is itself written in Tile Lang, showing the two projects are already working together in practice, not just in theory.
Is this actually a threat to Nvidia’s CUDA monopoly?
Not immediately, and not automatically. Reaching near the hardware ceiling on one chip is impressive, but an ecosystem is more than one optimized library. CUDA’s advantage comes from years of accumulated debuggers, profilers, documentation, tutorials, and existing production code that an enormous developer base already knows how to use. A new language or library only matters if developers actually pick it up and build on it, which takes years, not a single GitHub release.
What makes this release notable isn’t a grand claim that CUDA is finished. It’s the way it was released: MIT licensed, same API as the existing Nvidia-targeted library, built by the team that also builds the models that depend on this software, and published with working code rather than announcements. That’s a credible way to start building trust and adoption, even if it’s a long way from replacing an entrenched ecosystem.
Combined with Huawei’s parallel hardware push, including a super node system built from 128 Ascend 950 chips, the timing suggests a coordinated effort rather than an isolated stunt. Whether it actually chips away at CUDA’s grip depends on whether outside developers, not just DeepSeek and Huawei themselves, start building and shipping real projects on top of it.
Frequently Asked Questions
What is DeepGEMM Ascend?
It’s an open-source matrix multiplication library from DeepSeek, ported from their existing Nvidia-focused DeepGEMM library to run on Huawei’s Ascend AI chips, using the same API so existing code doesn’t need to be rewritten.
What is Tile Lang used for?
Tile Lang is a kernel-writing language that lets developers describe AI computations in Python-style code and compiles that down to fast, hardware-specific code. It already supports Nvidia, AMD, and Apple Metal, and now includes a native backend for Huawei’s Ascend 950 chip.
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
Does this mean Huawei chips now match Nvidia GPUs?
Not necessarily. The benchmarks show DeepGEMM Ascend uses nearly all of the Ascend chip’s own theoretical performance limit, which demonstrates efficient software. It doesn’t by itself show how that hardware ceiling compares to Nvidia’s newest GPUs.
Why is CUDA considered such a strong moat for Nvidia?
CUDA has roughly two decades of accumulated tooling, documentation, and an estimated 4 million developers worldwide who already know how to use it. That accumulated expertise and existing code is harder to displace than any single hardware performance gap.
Is DeepGEMM Ascend open source?
Yes. It was released on GitHub under an MIT license, with no accompanying keynote or major announcement, just the code and benchmark data for anyone to inspect or test independently.
