DeepSeek DeepGEMM: High-Performance GPU BLAS and MoE Kernel Library for SM90 and SM100
DeepSeek open-sources DeepGEMM, a clean GPU tensor core kernel library featuring runtime JIT compilation without pre-builds, FP8/FP4/BF16 GEMMs, communication-o
DeepSeek has open-sourced DeepGEMM, a high-performance GPU tensor core kernel library uniting the foundational computational primitives of modern large language models (LLMs) into a single, cohesive CUDA codebase.

Image source: deepseek-ai/DeepGEMM (GitHub)
DeepGEMM packages the critical building blocks required for modern LLM inference and serving—multi-precision GEMMs (FP8, FP4, BF16), fused Mixture of Experts with communication overlap (Mega MoE), MQA scoring for lightning indexers, and HyperConnection (HC). While drawing core concepts from NVIDIA CUTLASS and CuTe, it avoids heavy template algebra dependencies, providing concise and accessible kernel implementations for production use and learning.
Runtime JIT Compilation: Eliminating Pre-Build Overhead
Traditional high-performance CUDA libraries frequently require compiling extensive template permutations during installation, resulting in protracted build times and environment conflicts.
DeepGEMM circumvents this friction by incorporating a lightweight Just-In-Time (JIT) C++ compilation module. No CUDA pre-compilation is performed during initial package installation; instead, every kernel compiles dynamically at runtime based on the precise dimensions and active hardware architecture. This allows developers to integrate DeepGEMM directly into Python environments without build-phase friction.
Architecture Specialization for NVIDIA Hopper (SM90) and Blackwell (SM100)
DeepGEMM specifically targets current enterprise GPU architectures, offering fine-tuned execution paths for SM90 and SM100:
- Hopper (SM90): Provides fine-grained FP32 scaling factor support alongside optimized NT (Non-transposed / Transposed) memory layout GEMM execution.
- Blackwell (SM100): Introduces support for packed UE8M0 format and complete memory layout flexibility (NT, TN, NN, TT), alongside grouped GEMM dispatches designed to maximize tensor core occupancy.
- Multi-Precision Math: Delivers native control over FP8 and BF16 execution, as well as FP8xFP4 computation paths required by frontier model architectures.
Mega MoE: Fused Tensor Math and Overlapped NVLink Communication
A central component of DeepGEMM is its Mega MoE engine, architected to address bandwidth bottlenecks in distributed Mixture of Experts inference.
In conventional distributed MoE pipelines, expert routing communication and GPU tensor core computation run in separated phases, causing significant latency pauses. DeepGEMM resolves this by consolidating 11 previously separate kernels into a unified mega kernel while fusing shared experts directly into the execution path:
- 11-in-1 Kernel Consolidation and Shared Experts: Fuses routed expert computation and shared expert (SharedLinear 1/2) phases into a cohesive pipeline supporting both BF16 and FP8xFP4 execution.
- In-Kernel Dynamic Scheduling: Replaces host-side heuristics with an in-kernel task scheduler that interleaves L1 and L2 blocks, maximizing overlap between NVLink data transfers and tensor core execution waves.
- Quantifiable Speedups: Pull request benchmarks indicate this architecture runs Mega MoE 20% faster than previous iterations, translating into a 10%+ end-to-end inference speedup.
Hardware Requirements and Deployment Caveats
Before integrating DeepGEMM into model serving stacks, consider the following technical prerequisites and behavioral boundaries:
- Hardware and Software Prerequisites: Requires NVIDIA SM90 (Hopper) or SM100 (Blackwell) GPUs, Python 3.8+, PyTorch 2.1+, and a C++20-compliant compiler. For CUDA Toolkit, SM90 requires CUDA 12.3+ (12.9+ recommended for optimal performance), while SM100 requires CUDA 12.9+.
- No Automatic Pre-transformations: Pre-processing steps such as input tensor transposition or dynamic FP8 quantization casting are not automatically fused within DeepGEMM kernels. Because SM90 only supports the NT layout, upstream layers must supply inputs in the required shape and format.
- Symmetric Memory Allocations: Mega MoE relies on symmetric memory buffer slicing and an in-kernel dynamic task scheduler to support shared experts and overlap communication and computation across execution waves.
DeepGEMM offers clean, verifiable primitives that significantly reduce CUDA kernel complexity while squeezing peak efficiency out of modern datacenter GPUs.
Sources
- DeepSeek DeepGEMM GitHub Repository: deepseek-ai/DeepGEMM