Run a model locally
RunQwen/Qwen3-0.6B on your own computer and ask it a question. This small model is useful for trying the workflow; choose a larger model for more demanding tasks.
- NVIDIA GPU (Linux)
- Apple Silicon Mac
This example uses one CUDA 13-compatible NVIDIA GPU, Linux, Docker, and NVIDIA Container Toolkit. For setup, see Installation.Run this command on your computer. Docker pulls the image if needed, and SGLang downloads the model on first use:If you already installed SGLang with uv, run this in that Python environment instead:
The server is fired up and ready to roll!, open another terminal on the same computer and send a request:
choices[0].message.content. Both the model and the API server are running locally. Press Ctrl+C in the server terminal when you are done, or leave it running for the chat and batch examples below.
Serve Qwen3.8-27B
This configuration runsQwen/Qwen3.8-27B-FP8 on one NVIDIA H200. It requires Linux, a CUDA 13-compatible driver, Docker, and NVIDIA Container Toolkit configured for GPU access. For environment setup, see Installation. For other hardware configurations, see the Qwen3.8-27B Cookbook.
Set HF_CACHE_DIR to an absolute directory on the host for downloaded model weights, replacing the placeholder below. Run this command on the GPU host:
The server is fired up and ready to roll! before sending a request.
Already installed with uv?
Already installed with uv?
Run this command in the activated Python environment where you installed SGLang. It starts the same server without Docker:For Python environment setup, see Install with uv.
Qwen/Qwen3.8-27B-FP8 as the model name.
Connect to a model
The following Python examples connect to the local server above. You can also use a remote SGLang server; only the server needs hardware that can run the model. Install the Python client in your activated environment:localhost with its reachable address, use its model name, and supply an API key if authentication is enabled.
Stream a reply
Ask a question and print the answer as it arrives:Process multiple prompts
Send several independent requests concurrently, for example to summarize customer feedback:Build a customer support agent
Let an agent answer “Where is my order?” by looking up order data. The model chooses a tool, your application executes it, and the model uses the returned data to answer the customer. Use the Qwen3.8-27B server above, which enables a tool parser, or configure another model with the tool-calling guide. The small local example does not configure tool calling. This script uses the OpenAI client installed above; update the endpoint and model ID to match your server. Save this asorder_agent.py and run python order_agent.py. The order data is a local fixture, so no external service or API key is needed for the lookup.
lookup_order with an authorized order-system query to use real data.