Skip to main content
Choose an example for what you want to do:

Run a model locally

Run Qwen/Qwen3-0.6B on your own computer and ask it a question. This small model is useful for trying the workflow; choose a larger model for more demanding tasks.
This example uses one CUDA 13-compatible NVIDIA GPU, Linux, Docker, and NVIDIA Container Toolkit. For setup, see Installation.Run this command on your computer. Docker pulls the image if needed, and SGLang downloads the model on first use:
If you already installed SGLang with uv, run this in that Python environment instead:
Keep the server running. Once you see The server is fired up and ready to roll!, open another terminal on the same computer and send a request:
Look for the answer in choices[0].message.content. Both the model and the API server are running locally. Press Ctrl+C in the server terminal when you are done, or leave it running for the chat and batch examples below.

Serve Qwen3.8-27B

This configuration runs Qwen/Qwen3.8-27B-FP8 on one NVIDIA H200. It requires Linux, a CUDA 13-compatible driver, Docker, and NVIDIA Container Toolkit configured for GPU access. For environment setup, see Installation. For other hardware configurations, see the Qwen3.8-27B Cookbook. Set HF_CACHE_DIR to an absolute directory on the host for downloaded model weights, replacing the placeholder below. Run this command on the GPU host:
Docker downloads the image if needed. SGLang downloads the model on first use and starts the server. Keep this terminal open and wait for The server is fired up and ready to roll! before sending a request.
Run this command in the activated Python environment where you installed SGLang. It starts the same server without Docker:
For Python environment setup, see Install with uv.
Verify the running server with the Quickstart request, using Qwen/Qwen3.8-27B-FP8 as the model name.

Connect to a model

The following Python examples connect to the local server above. You can also use a remote SGLang server; only the server needs hardware that can run the model. Install the Python client in your activated environment:
Run this setup before each example in the same Python session:
These settings match the local Qwen3-0.6B server above. For a remote server, replace localhost with its reachable address, use its model name, and supply an API key if authentication is enabled.

Stream a reply

Ask a question and print the answer as it arrives:
The terminal displays the model’s explanation incrementally. You can use the same stream to update a chat interface. See the API guide for more request options.

Process multiple prompts

Send several independent requests concurrently, for example to summarize customer feedback:
You get one summary per input, in the original order. For batch processing inside a Python process without an HTTP server, see the offline engine API.

Build a customer support agent

Let an agent answer “Where is my order?” by looking up order data. The model chooses a tool, your application executes it, and the model uses the returned data to answer the customer. Use the Qwen3.8-27B server above, which enables a tool parser, or configure another model with the tool-calling guide. The small local example does not configure tool calling. This script uses the OpenAI client installed above; update the endpoint and model ID to match your server. Save this as order_agent.py and run python order_agent.py. The order data is a local fixture, so no external service or API key is needed for the lookup.
You should see the lookup result followed by an answer that the order has shipped and is expected on Friday. The wording can vary. SGLang serves the model; the Python application runs the tools and manages the conversation. Replace lookup_order with an authorized order-system query to use real data.