Skip to main content
SGLang Diffusion is SGLang’s built-in inference engine for image and video generation. It ships in the same repository and sglang Python package, using the shared sglang generate and sglang serve commands. It provides native SGLang pipelines, diffusers backend support, an OpenAI-compatible server, and an optimized kernel stack built on both precompiled sgl-kernel operators and JIT kernels for key inference paths.

Key Features

  • Broad model support across Wan, Hunyuan, Qwen-Image, FLUX, Z-Image, GLM-Image, and more
  • Fast inference with sgl-kernel, JIT kernels, scheduler improvements, and caching acceleration
  • Multiple interfaces: sglang generate, sglang serve, and an OpenAI-compatible API
  • Multi-platform support for NVIDIA, AMD, Intel XPU, Ascend, Apple Silicon, and Moore Threads

Quick Start

Start with the recommended Docker setup for Linux GPU deployments. The installation guide also covers pip/uv, source installation, and other platforms. Run the following commands inside the container or your activated Python environment. Generate an image:
Or start an HTTP server:

Start Here

Additional Documentation

Developer Documentation

References