sglang.multimodal_gen.benchmarks.bench_serving is a command-line tool designed to benchmark the online serving throughput and latency of diffusion models. It selects the image or video API from the requested task and offers flexible configurations for request rates, dataset types, and profiling.
1. Quick Start
1.1 Benchmarking in Low Concurrency
Run a benchmark on a local server (port 30000) generating 1 videos/images from thevbench dataset.
Command
1.2 Benchmarking in High Concurrency
Run a benchmark on a local server (port 30000) generating 20 videos/images from thevbench dataset.
Command
2. Parameter Reference
2.1 Connection Settings
| Argument | Default | Description |
|---|---|---|
--base-url | None | Base URL of the server (e.g., http://localhost:30000). If specified, this overrides --host and --port. |
--host | None | The server host (e.g., 127.0.0.1). |
--port | None | The server port. |
--model | None | Model name or path. |
2.2 Workload & Task Configuration
| Argument | Choices | Description |
|---|---|---|
--task | text-to-video, image-to-video, text-to-image, image-to-image, video-to-video | Defines the generation task when it cannot be inferred from the model metadata. |
--dataset | vbench, random | The source of prompts/inputs. |
--dataset-path | None | (Optional) Path to a local dataset file if not using built-in presets. |
--num-prompts | None | The total number of prompts/requests to execute during the benchmark. |
2.3 Generation Parameters
| Argument | Description |
|---|---|
--width | The target width for the generated image or video. |
--height | The target height for the generated image or video. |
--num-frames | Number of frames to generate (Specific to Video backends). |
--fps | Frames Per Second configuration (Specific to Video backends). |
2.4 Concurrency & Load Control
| Argument | Description |
|---|---|
--request-rate | The number of requests initiated per second. If set to inf, all requests are sent immediately (burst). If set to a number, request arrival times follow a Poisson process. |
--max-concurrency | The maximum number of requests allowed to execute simultaneously. This simulates a semaphore or upstream limit. Even if request-rate is high, the actual processing rate is capped by this value. |
2.5 Logging & Output
| Argument | Description |
|---|---|
--output-file | Path to save the benchmark metrics (JSON format). |
--disable-tqdm | If set, disables the progress bar in the console. |
3. Metrics
Request Throughput(req/s), Output Throughput (tok/s)Latency Mean(ms): Time to Per StepPeak Memory Max(ms): Max Memory Usage during running
