Skip to main content
This document provides commands for evaluating models’ accuracy and performance. Before open-sourcing new models, we strongly suggest running these commands to verify whether the score matches your internal benchmark results. For cross verification, please submit commands for installation, server launching, and benchmark running with all the scores and hardware requirements when open-sourcing your models. Reference: MiniMax M2

Accuracy

LLMs

SGLang provides built-in scripts to evaluate common benchmarks. MMLU
Command
GSM8K
Command
HellaSwag
Command
GPQA
Command
For reasoning models, add --thinking-mode <mode> (e.g., qwen3, deepseek-r1, deepseek-v3). You may skip it if the model has forced thinking enabled.
HumanEval
Command

VLMs

MMMU
Command
You can set max tokens by passing --extra-request-body '{"max_tokens": 4096}'.
For models capable of processing video, we recommend extending the evaluation to include VideoMME, MVBench, and other relevant benchmarks.

Performance

Performance benchmarks measure Latency (Time To First Token - TTFT) and Throughput (tokens/second).

LLMs

Latency-Sensitive Benchmark This simulates a scenario with low concurrency (e.g., single user) to measure latency.
Command
Throughput-Sensitive Benchmark This simulates a high-traffic scenario to measure maximum system throughput.
Command
Single Batch Performance You can also benchmark the performance of processing a single batch offline.
Command
You can run more granular benchmarks:
  • Low Concurrency: --num-prompts 10 --max-concurrency 1
  • Medium Concurrency: --num-prompts 80 --max-concurrency 16
  • High Concurrency: --num-prompts 500 --max-concurrency 100

Reporting Results

For each evaluation, please report:
  1. Metric Score: Accuracy % (LLMs and VLMs); Latency (ms) and Throughput (tok/s) (LLMs only).
  2. Environment settings: GPU type/count, SGLang commit hash.
  3. Launch configuration: Model path, TP size, and any special flags.
  4. Evaluation parameters: Number of shots, examples, max tokens.