--enable-metrics when you launch the server.
An example of the monitoring dashboard is available in examples/monitoring/grafana.json.
Language model metrics
sglang:time_to_first_token_seconds measures the time from request creation in the API server to the first token. For streaming requests it ends when the first output arrives after detokenization, which may contain no printable text. Non-streaming clients only see the final response, so for them it ends when the scheduler produced the first token.
Here is an example of the metrics:
Output
Request-level time per output token (TPOT)
sglang:request_time_per_output_token_seconds records TPOT once per successfully finished request:
is_streaming and reuses the --bucket-inter-token-latency buckets.
Unlike the token-weighted sglang:inter_token_latency_seconds, this metric gives every qualifying request equal weight.
p95 TPOT per model:
rate(..._sum[5m]) / rate(..._count[5m]). Per-request decode tokens/s is its reciprocal, so p5 TPS is 1 / p95 TPOT; the reciprocal of mean TPOT is not mean TPS.
Diffusion metrics
SGLang Diffusion exposes request, queue, stage and LoRA metrics with--enable-metrics. See Diffusion production metrics
for the metric reference, counting semantics and disaggregated scraping setup.
Setup Guide
This section describes how to set up the monitoring stack (Prometheus + Grafana) provided in theexamples/monitoring directory.
Prerequisites
- Docker and Docker Compose installed
- SGLang server running with metrics enabled
Usage
-
Start your SGLang server with metrics enabled:
ReplaceCommand
<your_model_path>with the actual path to your model (e.g.,meta-llama/Meta-Llama-3.1-8B-Instruct). Ensure the server is accessible from the monitoring stack (you might need--host 0.0.0.0if running in Docker). By default, the metrics endpoint will be available athttp://<sglang_server_host>:30000/metrics. -
Navigate to the monitoring example directory:
Command
-
Start the monitoring stack:
This command will start Prometheus and Grafana in the background.Command
-
Access the monitoring interfaces:
- Grafana: Open your web browser and go to http://localhost:3000.
- Prometheus: Open your web browser and go to http://localhost:9090.
-
Log in to Grafana:
- Default Username:
admin - Default Password:
adminYou will be prompted to change the password upon your first login.
- Default Username:
-
View the Dashboard:
The SGLang dashboard is pre-configured and should be available automatically. Navigate to
Dashboards->Browse->SGLang Monitoringfolder ->SGLang Dashboard.
Troubleshooting
- Port Conflicts: If you encounter errors like “port is already allocated,” check if other services (including previous instances of Prometheus/Grafana) are using ports
9090or3000. Usedocker psto find running containers anddocker stop <container_id>to stop them, or uselsof -i :<port>to find other processes using the ports. You might need to adjust the ports in thedocker-compose.yamlfile if they permanently conflict with other essential services on your system.
- Connection Issues:
- Ensure both Prometheus and Grafana containers are running (
docker ps). - Verify the Prometheus data source configuration in Grafana (usually auto-configured via
grafana/datasources/datasource.yaml). Go toConnections->Data sources->Prometheus. The URL should point to the Prometheus service (e.g.,http://prometheus:9090). - Confirm that your SGLang server is running and the metrics endpoint (
http://<sglang_server_host>:30000/metrics) is accessible from the Prometheus container. If SGLang is running on your host machine and Prometheus is in Docker, usehost.docker.internal(on Docker Desktop) or your machine’s network IP instead oflocalhostin theprometheus.yamlscrape configuration.
- Ensure both Prometheus and Grafana containers are running (
- No Data on Dashboard:
- Generate some traffic to your SGLang server to produce metrics. For example, run a benchmark:
Command
- Check the Prometheus UI (
http://localhost:9090) underStatus->Targetsto see if the SGLang endpoint is being scraped successfully. - Verify the
model_nameandinstancelabels in your Prometheus metrics match the variables used in the Grafana dashboard. You might need to adjust the Grafana dashboard variables or the labels in your Prometheus configuration.
- Generate some traffic to your SGLang server to produce metrics. For example, run a benchmark:
Configuration Files
The monitoring setup is defined by the following files within theexamples/monitoring directory:
docker-compose.yaml: Defines the Prometheus and Grafana services.prometheus.yaml: Prometheus configuration, including scrape targets.grafana/datasources/datasource.yaml: Configures the Prometheus data source for Grafana.grafana/dashboards/config/dashboard.yaml: Tells Grafana to load dashboards from the specified path.grafana/dashboards/json/sglang-dashboard.json: The actual Grafana dashboard definition in JSON format.
static_configs target in prometheus.yaml if your SGLang server runs on a different host or port.
Check if the metrics are being collected
Run:Output
Estimated Performance Metrics (MFU-related)
SGLang exports the following estimated per-GPU counters that can be used to derive Model FLOPs Utilization (MFU)-related signals:sglang:estimated_flops_per_gpu_total: Estimated floating-point operations.sglang:estimated_read_bytes_per_gpu_total: Estimated bytes read from memory.sglang:estimated_write_bytes_per_gpu_total: Estimated bytes written to memory.
--enable-metrics and
--enable-mfu-metrics are enabled.
These are cumulative counters. Use Prometheus rate(...) to get per-second values.
PromQL examples
Average TFLOPS per GPU:Notes
- These metrics are estimates intended for observability and trend analysis.
- Estimated memory bytes reflect modeled traffic and are not a direct hardware counter from GPU profilers.
