Skip to main content
SGLang provides Ollama API compatibility, allowing you to use the Ollama CLI and Python library with SGLang as the inference backend.

Prerequisites

You don’t need the Ollama server installed - SGLang acts as the backend. You only need the ollama CLI or Python library as the client.

Endpoints

EndpointMethodDescription
/GET, HEADHealth check for Ollama CLI
/api/tagsGETList available models
/api/chatPOSTChat completions (streaming & non-streaming)
/api/generatePOSTText generation (streaming & non-streaming)
/api/showPOSTModel information

Quick Start

1. Launch SGLang Server

The model name used with ollama run must match exactly what you passed to --model.

2. Use Ollama CLI

If connecting to a remote server behind a firewall:

3. Use Ollama Python Library

Example

Smart Router

For intelligent routing between local Ollama (fast) and remote SGLang (powerful) using an LLM judge, see the Smart Router documentation.

Summary

ComponentPurpose
Ollama APIFamiliar CLI/API that developers already know
SGLang BackendHigh-performance inference engine
Smart RouterIntelligent routing - fast local for simple tasks, powerful remote for complex tasks