- Output-preserving / lossless-style: system settings that should preserve model behavior while changing residency, parallelism, kernels, or scheduling.
- Quality-tradeoff / lossy or approximate: techniques that can change the denoising path, numerical representation, or generated output.
Start Here
- Pick a serving or generation mode from Deployment and Performance Modes.
--performance-mode autois the default; usespeedwhen the model fits in GPU memory and latency matters most,memorywhen GPU memory is the bottleneck, andmanualwhen every performance flag should be explicit. - Choose the right attention backend from Attention Backends.
- Use Sequence Parallelism only when the model and video shape benefit from sequence splitting.
- Use Inference Batching for concurrent compatible requests during serving.
- Use Profiling before changing several levers at once.
Choose a request quality tier
--quality is cumulative: a broader tier never drops an optimization from a
stricter tier.
Use
extra-high when you want to isolate fusion wins from approximate
acceleration. A tier may be a no-op when the active model has no eligible path.
Separately configured quantization, attention, or caching options still apply.
See Fused Kernels for the current
request-gated families and their numerical contracts.
Output-Preserving / Lossless-Style Levers
These settings should preserve model behavior while changing residency, parallelism, kernels, or scheduling. They are the first choices for production tuning.| Lever | Use when | Docs |
|---|---|---|
—performance-mode | You want a safe preset for speed or memory without overriding explicit flags. | Deployment and Performance Modes |
| Breakable CUDA graph | A supported pipeline serves a fixed set of shapes and eager execution is launch-bound. | CLI reference |
| Offload, FSDP, CFG parallelism | GPU memory, multi-GPU residency, or CFG branch splitting is the main bottleneck. | Deployment and Performance Modes |
| Sequence parallelism | Long image/video sequences need sequence-level parallelism. | Sequence Parallelism |
—encoder-parallel | Text/image encoding is a visible share of the request and the DiT replica sits idle during it. | Encoder Parallelism |
| Attention backend | Kernel choice dominates DiT latency or memory. | Attention Backends |
| Fused kernels | You want to know which elementwise chains are already fused, or to opt into the request-gated set. | Fused Kernels |
| Dynamic batching | Serving many compatible requests concurrently. | Inference Batching |
Quality-Tradeoff / Lossy Or Approximate Levers
These techniques can change the denoising path, numerical representation, or generated output. They are useful after you have a baseline and an acceptance criterion for quality.| Lever | Tradeoff | Docs |
|---|---|---|
| Cache-DiT | Skips selected DiT block or step computation based on cache decisions. | Cache-DiT |
| TeaCache | Reuses residuals when consecutive denoising steps are similar enough. | TeaCache |
| Progressive resolution | Runs early denoising at lower latent resolution for supported pipelines. | Progressive Resolution Generation |
| Quantization | Uses lower-precision transformer weights or activations. | Quantization |
Practical Order
- Establish a baseline with the target model, resolution, frame count, step count, and GPU type.
- Select
--performance-modeand explicit residency or parallelism flags. - Compare
quality=losslesswithquality=extra-highto isolate the request-gated fusion set. - Compare breakable CUDA graph against eager execution for supported fixed-shape pipelines. Pass every served resolution to
--warmup-resolutionsand confirm capture in the server log. Models with request-gated DiT fusions cannot combine those fusions with a graph captured from the lossless branches. - Tune attention backend and batching for the deployment pattern.
- Profile if the bottleneck is unclear.
- Add
quality=high, caching, progressive resolution, or quantization only after comparing output quality against your acceptance target.
