Performance & Runtime
Designed for low-latency, high-throughput inference with RadixAttention, prefix caching, and multi-GPU parallelism.
Models & Ecosystem
Broad support for Llama, Qwen, DeepSeek, and more. Compatible with Hugging
Face and OpenAI APIs.
Extensive Hardware Support
Native support across Hardware Platforms
including NVIDIA, AMD, Intel Xeon, Google TPU, Ascend NPU, and Moore Threads MUSA accelerators.
Community & Training
Open-source with widespread adoption, powering 400k+ GPUs and integrated with major RL frameworks.
Get Started
SGLang is an inference framework meant for production level serving. It is designed to deliver low-latency and high-throughput inference across a wide range of setups, from a single GPU to large distributed clusters.Quickstart
Start a model server and send a request to verify it is running.
Examples
Stream replies, request tool calls, and process multiple prompts from Python.
News and latest blogs
Community and resources
- Slack: Technical questions and development discussions.
- Events: Meetups, workshops, and office hours | Developer meetings.
- Updates: X | LinkedIn | LMSYS Blog.
- Contribute: Contributor guide | Roadmap.
- Release notes: Changes in each release.





