Skip to main content
SGLang integrates Cache-DiT, a caching acceleration engine for Diffusion Transformers (DiT), to achieve up to 1.69x inference speedup with minimal quality loss.

Overview

Cache-DiT uses intelligent caching strategies to skip redundant computation in the denoising loop:
  • DBCache (Dual Block Cache): Dynamically decides when to cache transformer blocks based on residual differences
  • TaylorSeer: Uses Taylor expansion for calibration to optimize caching decisions
  • DMD Calibrator: An exponential-basis forecasting calibrator (Dynamic Mode Decomposition, not Distribution Matching Distillation) that serves as a drop-in alternative to TaylorSeer’s polynomial basis; strongest on flow-matching models
  • SCM (Step Computation Masking): Step-level caching control for additional speedup
Cache-DiT also ships SVDQuant W4A4 (int4 / NVFP4) dynamic quantization, which can be combined with DBCache caching (see Quantization).

Basic Usage

Cache-DiT is a per-request switch: each request decides whether to run cached or lossless, and requests with different Cache-DiT settings never share a batch. The SGLANG_CACHE_DIT_* environment variables remain available as server-wide defaults for requests that leave the switch unset. Enable it for a single generation:
Or per request against a running server, via the OpenAI-compatible API:
enable_cache_dit accepts three states: true (on for this request), false (off for this request, overriding the server default), and unset (follow the SGLANG_CACHE_DIT_ENABLED server default). cache_dit_params accepts the DBCache knobs (Fn_compute_blocks, Bn_compute_blocks, max_warmup_steps, residual_diff_threshold, max_continuous_cached_steps, enable_taylorseer, taylorseer_order), the DMD knobs (enable_dmd, dmd_history, dmd_rank, dmd_ridge, dmd_svd_precision; DMD and TaylorSeer are mutually exclusive calibrators and cannot be enabled together), the SCM knobs (scm_preset, scm_compute_bins, scm_cache_bins, scm_policy), and a nested secondary dict with the DBCache knobs for the second transformer of dual-DiT models (unset secondary keys inherit the request’s primary values, then the SGLANG_CACHE_DIT_SECONDARY_* defaults). To make Cache-DiT the default for every request instead, export the environment variable when launching:

Diffusers Backend

Cache-DiT supports loading acceleration configs from a custom YAML file. For diffusers pipelines (diffusers backend), pass the YAML/JSON path via --cache-dit-config. This flow requires cache-dit >= 1.2.0 (cache_dit.load_configs).

Single GPU inference

Define a cache.yaml file that contains:
  • DBCache + TaylorSeer
Then apply the config with:
  • DBCache + TaylorSeer + SCM (Step Computation Mask)
Config
  • DBCache + TaylorSeer + SCM (Step Computation Mask) + Cache CFG
Config
  • DBCache + DMD Calibrator
Instead of TaylorSeer, you can use the DMD calibrator: an exponential-basis forecasting calibrator that serves as a drop-in alternative to TaylorSeer’s polynomial basis. DMD models the cached feature stream as a linear dynamical system (Y_{t+1} ~= A @ Y_t), forecasts cached features from the fitted eigen-modes, and stays accurate over longer cache skips where polynomial extrapolation diverges. DMD here refers to Dynamic Mode Decomposition (Schmid 2010), not Distribution Matching Distillation. DMD works best on flow-matching models (e.g., FLUX), while TaylorSeer is often better on DDPM-style models — try both. DMD and TaylorSeer are mutually exclusive — enable only one of enable_dmd / enable_taylorseer:
Config
A dmd_history window of 5–6 snapshots is typically the sweet spot — longer histories do not always help, because the feature dynamics drift across timesteps. With fewer than 4 uniformly spaced snapshots available, DMD transparently falls back to the Taylor expansion it maintains internally. See the Cache-DiT DMD documentation for the mathematical principle and quantitative comparisons. A ready-made config is available at examples/configs/cache_dmd.yaml in the Cache-DiT repository. Apply it with the same --cache-dit-config flag:

Distributed inference

  • 1D Parallelism
Define a parallelism only config yaml parallel.yaml file that contains:
Config
Then, apply the distributed inference acceleration config from yaml. ulysses_size: auto means that cache-dit will auto detect the world_size as the ulysses_size. Otherwise, you should manually set it as specific int number, e.g, 4. Then apply the distributed config with: (Note: please add --num-gpus N to specify the number of gpus for distributed inference)
  • 2D Parallelism
You can also define a 2D parallelism config yaml parallel_2d.yaml file that contains:
Config
Then, apply the 2D parallelism config from yaml. Here tp_size: 2 means using tensor parallelism with size 2. The ulysses_size: auto means that cache-dit will auto detect the world_size // tp_size as the ulysses_size.
  • 3D Parallelism
You can also define a 3D parallelism config yaml parallel_3d.yaml file that contains:
Config
Then, apply the 3D parallelism config from yaml. Here ulysses_size: 2, ring_size: 2, tp_size: 2 means using ulysses parallelism with size 2, ring parallelism with size 2 and tensor parallelism with size 2.
  • Ulysses Anything Attention
To enable Ulysses Anything Attention, you can define a parallelism config yaml parallel_uaa.yaml file that contains:
Config
  • Ulysses FP8 Communication
For device that don’t have NVLink support, you can enable Ulysses FP8 Communication to further reduce the communication overhead. You can define a parallelism config yaml parallel_fp8.yaml file that contains:
Config
  • Async Ulysses CP
You can also enable async ulysses CP to overlap the communication and computation. Define a parallelism config yaml parallel_async.yaml file that contains:
Config
Then, apply the config from yaml. Here ulysses_async: true means enabling async ulysses CP.
  • TE-P and VAE-P
You can also specify the extra parallel modules in the yaml config. For example, define a parallelism config yaml parallel_extra.yaml file that contains:
Config

Hybrid Cache and Parallelism

Define a hybrid cache and parallel acceleration config yaml hybrid.yaml file that contains:
Config
Then, apply the hybrid cache and parallel acceleration config from yaml.

Attention Backend

In some cases, users may want to only specify the attention backend without any other optimization configs. In this case, you can define a yaml file attention.yaml that only contains:
Config

Quantization

You can also specify the quantization config in the yaml file, required torchao>=0.16.0. For example, define a yaml file quantize.yaml that contains:
Config
Then, apply the quantization config from yaml. Please also enable torch.compile for better performance if you are using quantization. For example:
Command

SVDQuant (W4A4 int4 / NVFP4)

SVDQuant is Cache-DiT’s built-in W4A4 PTQ quantization (weights and activations in int4 or NVFP4, with smoothed low-rank branches). It can be freely combined with DBCache caching and the DMD calibrator for the largest speedups. ::::note SVDQuant requires a cache-dit build with CUDA extension support — a plain pip install cache-dit does NOT include it. Install one of:
Command
:::: Valid quant_type values are svdq_int4_r{32,64,128,256}_dq (int4 W4A4) and svdq_nvfp4_r{32,64,128,256}_dq (NVFP4 W4A4; requires a Blackwell GPU). Example config combining SVDQuant NVFP4 with DBCache + DMD (see examples/configs/blackwell/cache_dmd_svdq.yaml):
Config
For int4 W4A4 (pre-Blackwell GPUs), the same config with quant_type: "svdq_int4_r128_dq" is available at examples/configs/cache_dmd_svdq.yaml (add runtime_kernel: "v2" to svdq_kwargs). Enable torch.compile for the best SVDQuant performance, and make sure --warmup-steps covers the compile warmup (use the same value as --num-inference-steps):
Command
You can verify from the log that the quantization is active: [Cache-DiT] SVDQuant Type: svdq_nvfp4_r128_dq, Rank: 128.

Combined Configs: Cache + Parallelism + Quantization

You can also combine all the above configs together in a single yaml file combined.yaml that contains:
Config
Then, apply the combined cache, parallelism and quantization config from yaml. Please also enable torch.compile for better performance if you are using quantization.

Advanced Configuration

DBCache Parameters

DBCache controls block-level caching behavior:
ParameterEnv VariableDefaultDescription
FnSGLANG_CACHE_DIT_FN1Number of first blocks to always compute
BnSGLANG_CACHE_DIT_BN0Number of last blocks to always compute
WSGLANG_CACHE_DIT_WARMUP4Warmup steps before caching starts
RSGLANG_CACHE_DIT_RDT0.24Residual difference threshold
MCSGLANG_CACHE_DIT_MC3Maximum continuous cached steps

TaylorSeer Configuration

TaylorSeer improves caching accuracy using Taylor expansion:
ParameterEnv VariableDefaultDescription
EnableSGLANG_CACHE_DIT_TAYLORSEERfalseEnable TaylorSeer calibrator
OrderSGLANG_CACHE_DIT_TS_ORDER1Taylor expansion order (1 or 2)

DMD Calibrator Configuration

DMD (Dynamic Mode Decomposition, Schmid 2010 — not Distribution Matching Distillation) is an exponential-basis forecasting calibrator and a drop-in alternative to TaylorSeer’s polynomial basis. At each full-compute step it records a snapshot of the computed features; at a cached step it identifies a linear propagator from the recent snapshot window (one economy SVD with rank truncation, then eigendecomposition) and forecasts the current features via eigenvalue powers — cheap to advance, and stable over longer cache skips where polynomial extrapolation diverges. It typically improves both speed and quality over pure DBCache. DMD and TaylorSeer are mutually exclusive (enabling both raises a ValueError); DMD is best for flow-matching models, TaylorSeer for DDPM-style ones. See the Cache-DiT DMD documentation for details:
ParameterEnv VariableDefaultDescription
EnableSGLANG_CACHE_DIT_DMDfalseEnable the DMD calibrator
HistorySGLANG_CACHE_DIT_DMD_HISTORY6Snapshot window length; 5-6 typical. Needs >= 4 uniformly spaced snapshots, otherwise DMD falls back to TaylorSeer
RankSGLANG_CACHE_DIT_DMD_RANK0SVD truncation rank; 0 = automatic (drop modes below 1e-4 of the leading singular value)
RidgeSGLANG_CACHE_DIT_DMD_RIDGE1e-8Tikhonov regularization added to the inverted singular values
SVD PrecisionSGLANG_CACHE_DIT_DMD_SVD_PRECISIONmediumSVD precision: “low”, “medium” or “high”
Usage (SGLD backend, env-driven):
Command
On the diffusers backend, enable DMD from the yaml config instead (enable_dmd: true in cache_config, see Diffusers Backend); DMD can also be set per request via cache_dit_params: {"enable_dmd": true}.

Combined Configuration Example

DBCache and TaylorSeer are complementary strategies that work together, you can configure both sets of parameters simultaneously:
Command

SCM (Step Computation Masking)

SCM provides step-level caching control for additional speedup. It decides which denoising steps to compute fully and which to use cached results. SCM Presets SCM is configured with presets:
PresetCompute RatioSpeedQuality
none100%BaselineBest
slow~75%~1.3xHigh
medium~50%~2xGood
fast~35%~3xAcceptable
ultra~25%~4xLower
Usage
Command
Custom SCM Bins For fine-grained control over which steps to compute vs cache:
Command
SCM Policy
PolicyEnv VariableDescription
dynamicSGLANG_CACHE_DIT_SCM_POLICY=dynamicAdaptive caching based on content (default)
staticSGLANG_CACHE_DIT_SCM_POLICY=staticFixed caching pattern

Environment Variables

All Cache-DiT parameters can also be configured via environment variables, which act as the server-wide defaults for requests that don’t set enable_cache_dit / cache_dit_params. See Environment Variables for the complete list.

Supported Models

SGLang Diffusion x Cache-DiT supports almost all models originally supported in SGLang Diffusion:
Model FamilyExample Models
WanWan2.1, Wan2.2
FluxFLUX.1-dev, FLUX.2-dev
Z-ImageZ-Image-Turbo
QwenQwen-Image, Qwen-Image-Edit, Qwen-Image 2.1
HunyuanHunyuanVideo
MiniMaxMiniMax-H3 (T2VA, FL2VA, and Ref2VA)

Performance Tips

  1. Start with defaults: The default parameters work well for most models
  2. Use TaylorSeer: It typically improves both speed and quality
  3. Tune R threshold: Lower values = better quality, higher values = faster
  4. SCM for extra speed: Use medium preset for good speed/quality balance
  5. Warmup matters: Higher warmup = more stable caching decisions

Limitations

  • SGLang-native pipelines: Distributed Cache-DiT paths exist for supported pipelines. Hybrid SP+TP configurations add communication and cache coordination overhead, so validate them on the target model and hardware before using them as production defaults.
  • DiT layerwise offload: Compatible. Skipped blocks are not streamed, and the first layer after a skip may sync-load. Still incompatible with --use-fsdp-inference.
  • SCM minimum steps: SCM requires >= 8 inference steps to be effective. Some pipelines report steps - 1 NFEs (for example MiniMax-H3 num_inference_steps=8 is 7 NFEs), which trips the upstream steps_mask assertion; use at least 9 requested steps or custom bins.
  • Model support: The model must be registered in Cache-DiT’s BlockAdapterRegister or have an SGLang custom block adapter.

Troubleshooting

SCM disabled for low step count

For models with < 8 inference steps (e.g., DMD distilled models), SCM will be automatically disabled. DBCache acceleration still works.

SVDQuant unavailable or load failure

SVDQuant cases raise svdq_is_available() = False or undefined symbol: ... materialize_cow_storage ... when the installed cache-dit has no CUDA extension, or the prebuilt wheel was compiled against an incompatible torch. Fix: reinstall from the cache-dit-cu13 wheel matching your torch version, or build cache-dit from source with CACHE_DIT_BUILD_SVDQUANT=1 (see Quantization). Quick self-check:
Command

References