sglang/kernels/ops/diffusion.
This page is an inventory: what each kernel fuses, what its numerical contract is, and which models use it. It is not a lever you tune — most of these kernels are on by default and require no flag. The one switch is --quality, described below.
Numerical contracts and quality tiers
Multi-step denoising amplifies a per-step rounding difference into visible quality loss, so “close enough” and “bit-exact” are different products here. Thequality switch distinguishes unconditional bit-exact replacements from non-bit-exact eager-chain fusions:
Bit-exact — mounted unconditionally. The kernel reproduces every rounding boundary of the eager chain, so torch.equal holds against the reference. Some go quite far to get there: the fused LayerNorm+modulate kernel replicates PyTorch’s vectorized_layer_norm_kernel down to its Welford update order, guarded reciprocal, and warp-fold tree; the fused RMSNorm+scale/shift kernel replicates FlashInfer’s CuTe-DSL RMSNormKernel fragment order and shfl.bfly fold. Because the dispatch they replicate can change underneath them, each one still verifies itself against the live eager chain on first sight and falls back permanently on any mismatch.
Not bit-exact — request-gated. These differ from eager only at half-precision rounding-order level, but that is enough to matter, so they are mounted only for quality="extra-high" and quality="high" requests, at batch boundaries, all-or-nothing per transformer. The default quality="lossless" runs the unmodified reference chain.
Model/checkpoint-native. Generic close-contract kernels, sparse operators, and FP8/NVFP4 producers can be part of a model implementation or a separately selected deployment path. They are documented in the inventory, but quality does not select or undo those choices.
Selection-equivalent routing — enabled unconditionally. LingBot Video’s fused group-limited top-k returns the same selected expert-id set as its guarded CUDA torch.topk(..., sorted=False) reference chain. The order of those ids is not part of either path’s contract. Because the selected experts are unchanged, this path does not depend on the request quality tier.
A plain fp32 single-pass norm fusion looks harmless and is not. On ERNIE-Image it moved the 50-step trajectory to 18.83 dB PSNR, which is what motivated the bit-exact rewrite of that path.
If a model has no eligible request-gated fusion,
extra-high can execute the same path as lossless. Likewise, high adds only the model-specific high-only paths that the active pipeline implements.
quality is not a master precision switch. A quantized checkpoint, an explicitly selected approximate attention backend, or an independently enabled cache remains active at every quality tier.Enabling the request-gated set
lossless; the OpenAI-compatible endpoints carry it per request. Images:
quality participates in the dynamic-batch signature, so mixed-quality traffic is batched separately and the transition happens safely at a batch boundary. Mounting is all-or-nothing: if any marked site on a transformer fails its static guards, no site on that transformer is fused.
These fusion families mount under both quality="extra-high" and
quality="high":
Kernel inventory
45 operators are registered in the kernel registry across 51 implementations (some operators carry several backends). Backends are named by provenance, not device:KDA identifies Kernel Design Agents implementations, JIT compiles under nvcc and hipcc, TRITON identifies Triton sources, CUTE_DSL needs CUTLASS, FLYDSL is ROCm gfx950 only, and AOT comes from the sgl_kernel wheel. Per-operator capability metadata determines which devices can load each implementation.
Normalization
adaLN modulation and gating
LingBot World camera conditioning also uses
modulate_scale_shift with one
affine row per token. The causal normalization and residual paths verify each
new signature outside CUDA graph capture and fall back to their native chains
on mismatch or unsupported layouts. These paths preserve lossless arithmetic;
kernel replay support does not enable BCG for a model configuration that disables it.
RoPE and QK-norm
The Klein strided path preserves the BF16 rounding between RMSNorm and RoPE,
including signed zeros. It reads packed Q/K views without changing the adjacent
V or MLP values. Unsupported layouts, devices, or RMSNorm dispatches use the
original helper. A new signature is verified outside CUDA Graph capture; an
unverified capture falls back.
SGLANG_ENABLE_FUSED_QKNORM_ROPE=0 also disables
this path. This kernel does not enable model-level BCG support for FLUX.2.
Activation
Attention
MoE routing
Data movement
Every kernel here only moves values (plus zero fill, plus at most one same-order add), so each is bitwise identical to the aten chain it replaces.
Joy Image Edit uses the joint-copy path for eligible CUDA FP16/BF16 image
tensors of at least 32 MiB. It preserves image-first token order, copies
values without arithmetic, and verifies each new shape/stride signature
against the native concatenations before enabling it. Small inputs,
unsupported layouts, gradient-bearing inputs, and unverified signatures
during graph capture use the native path. This does not enable model-level
BCG support.
For large Hopper BF16 image Q/K with 32 heads of width 128, Joy also uses
the existing out-of-place QK-Norm + RoPE kernel to read the packed projection
directly. This removes two input copies while retaining the original CUDA
arithmetic and contiguous outputs. Each new signature is checked bitwise;
inputs remain intact if the operation fails. Other shapes and platforms,
compilation, unverified capture, and
SGLANG_ENABLE_FUSED_QKNORM_ROPE=0
retain the original helper.
Quantized layout producers
These kernels preserve the quantized checkpoint path’s selected reference operation. They are not a claim that FP8 or NVFP4 is equivalent to an unquantized BF16 checkpoint.Coverage by model
Kernels are written against a specific eager chain in a specific model, so coverage is per-model rather than universal.Inspecting what is registered
Every kernel is described by aKernelSpec in the process-wide registry, so the inventory is queryable without importing any backend:
Importing the kernels
Runtime code imports from the package, never from a submodule:can_use_<op>(...) first and fall back to the reference chain when it returns False; the kernel raises on an unsupported input rather than silently returning None.
The package README.md carries a selection matrix for the cases where several kernels look interchangeable and are not. The normalization domain alone holds more than a dozen implementations that differ by numerical contract, activation layout, and backend rather than by speed.
References
- Performance Optimization
- Attention Backends
- Quantization
- Profiling
sglang/kernels/ops/diffusion— source and selection matrix- RFC #29630 — the unified
sglang.kernelsnamespace
