sglang Python package, using the shared sglang generate and sglang serve commands.
It provides native SGLang pipelines, diffusers backend support, an OpenAI-compatible server, and an optimized kernel stack built on both precompiled sgl-kernel operators and JIT kernels for key inference paths.
Key Features
- Broad model support across Wan, Hunyuan, Qwen-Image, FLUX, Z-Image, GLM-Image, and more
- Fast inference with
sgl-kernel, JIT kernels, scheduler improvements, and caching acceleration - Multiple interfaces:
sglang generate,sglang serve, and an OpenAI-compatible API - Multi-platform support for NVIDIA, AMD, Intel XPU, Ascend, Apple Silicon, and Moore Threads
Quick Start
Start with the recommended Docker setup for Linux GPU deployments. The installation guide also covers pip/uv, source installation, and other platforms. Run the following commands inside the container or your activated Python environment. Generate an image:Start Here
- Installation: install SGLang Diffusion and platform dependencies
- Supported Models: browse supported model families, tasks, and public checkpoints
- CLI: run one-off generation jobs or launch a persistent server
- OpenAI-Compatible API: send image and video requests to the HTTP server
- ComfyUI plugin: use SGLang from ComfyUI in server mode or as a per-step DiT backend
- Performance Overview: choose speed, memory, parallelism, caching, and quality-tradeoff levers
- Caching Acceleration: use Cache-DiT, TeaCache, or Spectrum to reduce denoising cost
- Quantization: configure component checkpoint and causal KV-cache quantization
- Realtime and Causal Video Models: understand session state, causal caches, and realtime-only controls
- Contributing: contribution workflow, adding new models, and CI perf baselines
Additional Documentation
- Post-Processing: frame interpolation and upscaling
- Models with AR Stage: run hybrid diffusion pipelines like GLM-Image with a separate AR encoder server
- Models with Prompt Enhancement: run ERNIE-Image with either built-in PE or a separate PE server
- Deployment and Performance Modes: choose
--performance-mode, offload, FSDP, CFG parallelism, SP, and TP - Attention Backends: choose the best backend for your model and hardware
- Sequence Parallelism: configure SP, Ulysses, and ring-based splitting for long sequences
- Encoder Parallelism: fold, data-parallel, or replicate the text/image encoders across idle GPUs
- Inference Batching: batch compatible native diffusion requests during serving
- Progressive Resolution Generation: run early denoising steps at lower latent resolution for selected pipelines
- Environment Variables: platform, caching, storage, and debugging configuration
Developer Documentation
- Support New Models: implementation guide for new diffusion pipelines
- CI Performance Baselines: generate and update performance baselines used in CI
