Skip to main content
Star Fork

Performance & Runtime

Designed for low-latency, high-throughput inference with RadixAttention, prefix caching, and multi-GPU parallelism.

Models & Ecosystem

Broad support for Llama, Qwen, DeepSeek, and more. Compatible with Hugging Face and OpenAI APIs.

Extensive Hardware Support

Native support across Hardware Platforms including NVIDIA, AMD, Intel Xeon, Google TPU, Ascend NPU, and Moore Threads MUSA accelerators.

Community & Training

Open-source with widespread adoption, powering 400k+ GPUs and integrated with major RL frameworks.
SGLang powers large-scale production deployments, generating trillions of tokens each day across more than 400,000 GPUs worldwide. It is hosted under the non-profit open-source organization LMSYS.

Get Started

SGLang is an inference framework meant for production level serving. It is designed to deliver low-latency and high-throughput inference across a wide range of setups, from a single GPU to large distributed clusters.

Quickstart

Start a model server and send a request to verify it is running.

Examples

Stream replies, request tool calls, and process multiple prompts from Python.
For environment setup and other installation methods, see the Installation guide.

News and latest blogs

Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache
SGLang and Miles Add Day-0 Support for DeepSeek-V4.1
Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack
Infer-forge: Harness, Loop, and Graph Engineering Around SGLang
MiniMax-H3 on 8\u00d7H200: 1.95\u00d7 Lossless, Up to 6.24\u00d7 at 0.76\u20130.91 SSIM
Qwen3.8-Flash-Next: Day-0 Support in SGLang

Community and resources