Overview
Cache-DiT uses intelligent caching strategies to skip redundant computation in the denoising loop:- DBCache (Dual Block Cache): Dynamically decides when to cache transformer blocks based on residual differences
- TaylorSeer: Uses Taylor expansion for calibration to optimize caching decisions
- DMD Calibrator: An exponential-basis forecasting calibrator (Dynamic Mode Decomposition, not Distribution Matching Distillation) that serves as a drop-in alternative to TaylorSeer’s polynomial basis; strongest on flow-matching models
- SCM (Step Computation Masking): Step-level caching control for additional speedup
Basic Usage
Cache-DiT is a per-request switch: each request decides whether to run cached or lossless, and requests with different Cache-DiT settings never share a batch. TheSGLANG_CACHE_DIT_* environment variables remain available as
server-wide defaults for requests that leave the switch unset.
Enable it for a single generation:
enable_cache_dit accepts three states: true (on for this request), false
(off for this request, overriding the server default), and unset (follow the
SGLANG_CACHE_DIT_ENABLED server default). cache_dit_params accepts the
DBCache knobs (Fn_compute_blocks, Bn_compute_blocks, max_warmup_steps,
residual_diff_threshold, max_continuous_cached_steps, enable_taylorseer,
taylorseer_order), the DMD knobs (enable_dmd, dmd_history, dmd_rank,
dmd_ridge, dmd_svd_precision; DMD and TaylorSeer are mutually exclusive
calibrators and cannot be enabled together), the SCM knobs (scm_preset,
scm_compute_bins, scm_cache_bins, scm_policy), and a nested secondary
dict with the DBCache
knobs for the second transformer of dual-DiT models (unset secondary keys
inherit the request’s primary values, then the
SGLANG_CACHE_DIT_SECONDARY_* defaults).
To make Cache-DiT the default for every request instead, export the
environment variable when launching:
Diffusers Backend
Cache-DiT supports loading acceleration configs from a custom YAML file. For diffusers pipelines (diffusers backend), pass the YAML/JSON path via --cache-dit-config. This
flow requires cache-dit >= 1.2.0 (cache_dit.load_configs).
Single GPU inference
Define acache.yaml file that contains:
- DBCache + TaylorSeer
- DBCache + TaylorSeer + SCM (Step Computation Mask)
Config
- DBCache + TaylorSeer + SCM (Step Computation Mask) + Cache CFG
Config
- DBCache + DMD Calibrator
Y_{t+1} ~= A @ Y_t), forecasts
cached features from the fitted eigen-modes, and stays accurate over longer cache skips where
polynomial extrapolation diverges. DMD here refers to Dynamic Mode Decomposition (Schmid
2010), not Distribution Matching Distillation. DMD works best on flow-matching models
(e.g., FLUX), while TaylorSeer is often better on DDPM-style models — try both. DMD and
TaylorSeer are mutually exclusive — enable only one of enable_dmd / enable_taylorseer:
Config
dmd_history window of 5–6 snapshots is typically the sweet spot — longer histories do not
always help, because the feature dynamics drift across timesteps. With fewer than 4 uniformly
spaced snapshots available, DMD transparently falls back to the Taylor expansion it maintains
internally. See the
Cache-DiT DMD documentation
for the mathematical principle and quantitative comparisons. A ready-made config is available
at
examples/configs/cache_dmd.yaml
in the Cache-DiT repository. Apply it with the same --cache-dit-config flag:
Distributed inference
- 1D Parallelism
parallel.yaml file that contains:
Config
ulysses_size: auto means that cache-dit will auto detect the world_size as the ulysses_size. Otherwise, you should manually set it as specific int number, e.g, 4.
Then apply the distributed config with: (Note: please add --num-gpus N to specify the number of gpus for distributed inference)
- 2D Parallelism
parallel_2d.yaml file that contains:
Config
tp_size: 2 means using tensor parallelism with size 2. The ulysses_size: auto means that cache-dit will auto detect the world_size // tp_size as the ulysses_size.
- 3D Parallelism
parallel_3d.yaml file that contains:
Config
ulysses_size: 2, ring_size: 2, tp_size: 2 means using ulysses parallelism with size 2, ring parallelism with size 2 and tensor parallelism with size 2.
- Ulysses Anything Attention
parallel_uaa.yaml file that contains:
Config
- Ulysses FP8 Communication
parallel_fp8.yaml file that contains:
Config
- Async Ulysses CP
parallel_async.yaml file that contains:
Config
ulysses_async: true means enabling async ulysses CP.
- TE-P and VAE-P
parallel_extra.yaml file that contains:
Config
Hybrid Cache and Parallelism
Define a hybrid cache and parallel acceleration config yamlhybrid.yaml file that contains:
Config
Attention Backend
In some cases, users may want to only specify the attention backend without any other optimization configs. In this case, you can define a yaml fileattention.yaml that only contains:
Config
Quantization
You can also specify the quantization config in the yaml file, requiredtorchao>=0.16.0. For example, define a yaml file quantize.yaml that contains:
Config
Command
SVDQuant (W4A4 int4 / NVFP4)
SVDQuant is Cache-DiT’s built-in W4A4 PTQ quantization (weights and activations in int4 or NVFP4, with smoothed low-rank branches). It can be freely combined with DBCache caching and the DMD calibrator for the largest speedups. ::::note SVDQuant requires a cache-dit build with CUDA extension support — a plainpip install cache-dit does NOT include it. Install one of:
Command
quant_type values are svdq_int4_r{32,64,128,256}_dq (int4 W4A4) and
svdq_nvfp4_r{32,64,128,256}_dq (NVFP4 W4A4; requires a Blackwell GPU). Example config
combining SVDQuant NVFP4 with DBCache + DMD (see
examples/configs/blackwell/cache_dmd_svdq.yaml):
Config
quant_type: "svdq_int4_r128_dq" is available at
examples/configs/cache_dmd_svdq.yaml
(add runtime_kernel: "v2" to svdq_kwargs).
Enable torch.compile for the best SVDQuant performance, and make sure --warmup-steps
covers the compile warmup (use the same value as --num-inference-steps):
Command
[Cache-DiT] SVDQuant Type: svdq_nvfp4_r128_dq, Rank: 128.
Combined Configs: Cache + Parallelism + Quantization
You can also combine all the above configs together in a single yaml filecombined.yaml that contains:
Config
Advanced Configuration
DBCache Parameters
DBCache controls block-level caching behavior:| Parameter | Env Variable | Default | Description |
|---|---|---|---|
| Fn | SGLANG_CACHE_DIT_FN | 1 | Number of first blocks to always compute |
| Bn | SGLANG_CACHE_DIT_BN | 0 | Number of last blocks to always compute |
| W | SGLANG_CACHE_DIT_WARMUP | 4 | Warmup steps before caching starts |
| R | SGLANG_CACHE_DIT_RDT | 0.24 | Residual difference threshold |
| MC | SGLANG_CACHE_DIT_MC | 3 | Maximum continuous cached steps |
TaylorSeer Configuration
TaylorSeer improves caching accuracy using Taylor expansion:| Parameter | Env Variable | Default | Description |
|---|---|---|---|
| Enable | SGLANG_CACHE_DIT_TAYLORSEER | false | Enable TaylorSeer calibrator |
| Order | SGLANG_CACHE_DIT_TS_ORDER | 1 | Taylor expansion order (1 or 2) |
DMD Calibrator Configuration
DMD (Dynamic Mode Decomposition, Schmid 2010 — not Distribution Matching Distillation) is an exponential-basis forecasting calibrator and a drop-in alternative to TaylorSeer’s polynomial basis. At each full-compute step it records a snapshot of the computed features; at a cached step it identifies a linear propagator from the recent snapshot window (one economy SVD with rank truncation, then eigendecomposition) and forecasts the current features via eigenvalue powers — cheap to advance, and stable over longer cache skips where polynomial extrapolation diverges. It typically improves both speed and quality over pure DBCache. DMD and TaylorSeer are mutually exclusive (enabling both raises aValueError); DMD is
best for flow-matching models, TaylorSeer for DDPM-style ones. See the
Cache-DiT DMD documentation
for details:
| Parameter | Env Variable | Default | Description |
|---|---|---|---|
| Enable | SGLANG_CACHE_DIT_DMD | false | Enable the DMD calibrator |
| History | SGLANG_CACHE_DIT_DMD_HISTORY | 6 | Snapshot window length; 5-6 typical. Needs >= 4 uniformly spaced snapshots, otherwise DMD falls back to TaylorSeer |
| Rank | SGLANG_CACHE_DIT_DMD_RANK | 0 | SVD truncation rank; 0 = automatic (drop modes below 1e-4 of the leading singular value) |
| Ridge | SGLANG_CACHE_DIT_DMD_RIDGE | 1e-8 | Tikhonov regularization added to the inverted singular values |
| SVD Precision | SGLANG_CACHE_DIT_DMD_SVD_PRECISION | medium | SVD precision: “low”, “medium” or “high” |
Command
enable_dmd: true in cache_config, see
Diffusers Backend); DMD can also be set per request via
cache_dit_params: {"enable_dmd": true}.
Combined Configuration Example
DBCache and TaylorSeer are complementary strategies that work together, you can configure both sets of parameters simultaneously:Command
SCM (Step Computation Masking)
SCM provides step-level caching control for additional speedup. It decides which denoising steps to compute fully and which to use cached results. SCM Presets SCM is configured with presets:| Preset | Compute Ratio | Speed | Quality |
|---|---|---|---|
none | 100% | Baseline | Best |
slow | ~75% | ~1.3x | High |
medium | ~50% | ~2x | Good |
fast | ~35% | ~3x | Acceptable |
ultra | ~25% | ~4x | Lower |
Command
Command
| Policy | Env Variable | Description |
|---|---|---|
dynamic | SGLANG_CACHE_DIT_SCM_POLICY=dynamic | Adaptive caching based on content (default) |
static | SGLANG_CACHE_DIT_SCM_POLICY=static | Fixed caching pattern |
Environment Variables
All Cache-DiT parameters can also be configured via environment variables, which act as the server-wide defaults for requests that don’t setenable_cache_dit / cache_dit_params.
See Environment Variables for the complete list.
Supported Models
SGLang Diffusion x Cache-DiT supports almost all models originally supported in SGLang Diffusion:| Model Family | Example Models |
|---|---|
| Wan | Wan2.1, Wan2.2 |
| Flux | FLUX.1-dev, FLUX.2-dev |
| Z-Image | Z-Image-Turbo |
| Qwen | Qwen-Image, Qwen-Image-Edit, Qwen-Image 2.1 |
| Hunyuan | HunyuanVideo |
| MiniMax | MiniMax-H3 (T2VA, FL2VA, and Ref2VA) |
Performance Tips
- Start with defaults: The default parameters work well for most models
- Use TaylorSeer: It typically improves both speed and quality
- Tune R threshold: Lower values = better quality, higher values = faster
- SCM for extra speed: Use
mediumpreset for good speed/quality balance - Warmup matters: Higher warmup = more stable caching decisions
Limitations
- SGLang-native pipelines: Distributed Cache-DiT paths exist for supported pipelines. Hybrid SP+TP configurations add communication and cache coordination overhead, so validate them on the target model and hardware before using them as production defaults.
- DiT layerwise offload: Compatible. Skipped blocks are not streamed, and the first layer after a skip may sync-load. Still incompatible with
--use-fsdp-inference. - SCM minimum steps: SCM requires >= 8 inference steps to be effective. Some pipelines report
steps - 1NFEs (for example MiniMax-H3num_inference_steps=8is 7 NFEs), which trips the upstreamsteps_maskassertion; use at least 9 requested steps or custom bins. - Model support: The model must be registered in Cache-DiT’s
BlockAdapterRegisteror have an SGLang custom block adapter.
Troubleshooting
SCM disabled for low step count
For models with < 8 inference steps (e.g., DMD distilled models), SCM will be automatically disabled. DBCache acceleration still works.SVDQuant unavailable or load failure
SVDQuant cases raisesvdq_is_available() = False or
undefined symbol: ... materialize_cow_storage ... when the installed cache-dit has no CUDA
extension, or the prebuilt wheel was compiled against an incompatible torch. Fix: reinstall
from the cache-dit-cu13 wheel matching your torch version, or build cache-dit from source
with CACHE_DIT_BUILD_SVDQUANT=1 (see Quantization). Quick self-check:
Command
