Skip to main content

Deployment

For all methods and hardware platforms, see the official SGLang installation guide. The two paths below match the Python / Docker toggle in the command panel. PPLX-Decider support landed after v0.5.21, so until a release includes it, install a nightly build.
Command
Then run the Python output of the command panel below in that environment.
Pick your GPU to generate the launch command. The model runs on one GPU in BF16, and the decision support needs no flag: SGLang reads the checkpoint’s decision_config.json, loads its decision readout, and answers on /v1/systemone with the prompt, answer codes, and calibrated temperature the model was trained with. The GB300 recipe adds --disable-radix-cache for speed, as explained in Configuration Tips.
The H200 recipe is the original one, validated against the checkpoint’s own reference implementation on one H200 when support landed (#42183), and it has no speed numbers yet. The GB300 recipe adds --disable-radix-cache (see Configuration Tips) and was measured end to end on one GB300 at SGLang commit 70f0b7351e: four request shapes from 1 to 256 concurrent requests, Belebele and WinoGrande accuracy, and per-question parity with the reference implementation. §3 has the method and the full tables.

Playground

The Playground layers SGLang features on top of the recipe above. Any override changes the badge to Not Verified until that exact configuration is tested end to end. The Configuration Tips say which overrides were measured on GB300 and what they did. Tensor parallelism times data-parallel replicas must fit the GPUs on one node.

1. Model Introduction

PPLX-Decider-v1-27B is a decision model from Perplexity, fine-tuned from Qwen3.8-27B under the Apache-2.0 license. It replaces the language model head with a readout over 255 answer codes and answers choice, yes or no (noul), and score questions about a text state and optional images with a probability for every option, without generating text. Each question is one prefill pass, so the cost of a decision is the cost of reading its prompt. Its successor, PPLX-Decider-v1.1-27B, adds noncausal full attention and more training data. The backbone is the Qwen3.8-27B hybrid: 64 layers in which 48 Gated DeltaNet (linear attention) layers alternate with 16 gated full-attention layers, a 5120-wide hidden state, and the Qwen3.8 vision encoder, for 26.1B parameters (52 GB in BF16). The checkpoint ships a bare Qwen3_5Model backbone, the 255-row readout in readout.safetensors, and decision_config.json with the answer-code token ids and a calibrated temperature of 2.2076. SGLang serves it as Qwen3_5ForConditionalGeneration and writes the readout into the language model head rows of the code tokens, so the probabilities come out of the regular scoring path.
PropertyValue
Checkpointperplexity-ai/pplx-decider-v1-27b (BF16)
Base modelQwen3.8-27B, one epoch of supervised fine-tuning on decision data
Question typeschoice (up to 255 options), noul (yes or no), score (up to 10 levels)
InputsText or JSON state, optional images
Prompt lengthTrained on question prompts of up to 8,192 tokens, and the backbone accepts 262,144
LicenseApache-2.0
Perplexity reports 85.71% overall accuracy across 11 public benchmarks for this checkpoint, against 74.76% for Qwen3.8-27B, measured through the Perplexity API. The per-benchmark table is on the model card. §3 measures accuracy and speed on SGLang. Resources: HuggingFace, reference inference code, Decision models in SGLang.

2. Configuration Tips

  • Route. Send requests to /v1/systemone, the System One API the checkpoint was trained for. /v1/decisions refuses this checkpoint because its labels and prompt are not the ones the readout learned, and chat or generation requests are not meaningful without a language model head.
  • Clients. Clients of the System One API, including the TypeSafe SDKs, work by pointing their base URL at the server. See System One compatible API for the request and response reference.
  • Images. Pass each image in images as base64 bytes, a base64 data URL, or an http(s) URL. They precede the text in every question, as in the checkpoint’s reference code. The checkpoint’s processor resizes every image to between 65,536 and 262,144 pixels, so one image costs 64 to 256 prompt tokens: a 1024×768 screenshot is 234 tokens, and a 4K one is no more expensive than 1080p (252). Crop to the region that matters rather than sending a larger image.
  • What a decision costs. Each question is its own prompt: an ~84-token wrapper (system message, chat template, answer instructions), the state, the question, and its options, read in one prefill pass. Several questions in one request are scored together, but each re-reads the whole state, so the cost of a request is roughly questions × state tokens.
  • Prefix cache. On this hybrid model the prefix cache resumes from where an earlier prompt ended, because that is where the Gated DeltaNet state is saved, not from any shared prefix. An identical request measured 215 ms cold and 56 ms cached on GB300, but a different question about the same state, and the other questions in the same request, reused none of the state. The GB300 recipe therefore turns the cache off with --disable-radix-cache: unique-state traffic loses no reuse, each request holds one state slot instead of five, and the per-request state bookkeeping goes away, which took a short question from 59 to 51 ms and an image request from 126 to 87 ms with the attention kernel held fixed. Turn it back on only if your traffic repeats identical requests, the same state and the same question, often enough to matter.
  • Prompt length. The model was fine-tuned on question prompts of up to 8,192 tokens, and the reference code refuses longer ones. SGLang serves the backbone’s full 262,144-token context, so keep states within the trained length. A longer prompt is answered, but its accuracy is untested.
  • Throughput is compute-bound. A decision is pure prefill, and one GB300 reads about 27,000 prompt tokens per second in BF16 once batches are full: about 71 requests per second for a 382-token question, and 3.3 for an 8K-token state. Past that point, more concurrency only adds queueing. The tables in §3 map concurrency to latency, so pick the highest concurrency that fits your latency budget.
  • Memory. The weights take 52 GB, and the default --mem-fraction-static hands the rest of the GPU to the KV and Gated DeltaNet state pools. The startup log line max_running_requests is capped to ... by the mamba state cache is expected: on GB300, each request reserves five state slots with the prefix cache on (a cap of 118) and one with it off (593). Even the lower cap only binds for questions shorter than about 140 tokens, because a prefill batch fills its 16,384-token chunk first. --mamba-ssm-dtype bfloat16 halves the state pool but measured no faster, so keep the checkpoint’s FP32 state.
  • GB300 kernels. With the prefix cache off, SGLang picks the TRT-LLM attention kernel with 64-token pages on GB300 (Triton with 1-token pages while the cache is on), and that kernel is where the long-state gain comes from: an 8K-token state drops from 377 to 325 ms and saturates at 3.30 instead of 2.87 requests per second, the same as forcing --attention-backend trtllm_mha --page-size 64 with the cache on. FlashInfer’s Gated DeltaNet prefill (--linear-attn-prefill-backend flashinfer) adds another 4-8% on text, taking the 8K-token state to 300 ms and 3.56 requests per second, but it made image requests slower (101 vs 87 ms alone, 9% less throughput), so the recipe leaves it to the Playground for text-only traffic. A larger prefill chunk (--chunked-prefill-size 32768) did not help: one 16,384-token chunk already keeps the GPU busy.
  • More GPUs. The model fits one GPU, so add throughput with data-parallel replicas: --dp-size 2 on two GB300s served 139 short questions per second (1.96× one GPU) and 6.5 8K-token states per second, and because the questions of one request are spread over the replicas, it also cut four questions on a 1K-token state from 179 to 104 ms. The extra dispatch hop added about 15 ms to a lone short question. Tensor parallelism only pays off for latency on long states: --tp 2 took an 8K-token state from 325 to 208 ms, but delivered less throughput than two replicas and made a short question slower (68 ms), because the all-reduce dominates small batches. SGLang Model Gateway does not route /v1/systemone yet, so to spread load over several servers instead of --dp-size, use a plain HTTP load balancer.
  • FP8. --quantization fp8 quantizes the BF16 weights at load time and measured about 20% more throughput on GB300 (84 vs 70 requests per second for short questions), with the same Belebele and WinoGrande scores. But its probabilities moved by 0.010 on average and up to 0.20 against the reference implementation, four times the drift of any BF16 configuration, and 0.9% of answers changed their top option. The calibrated probabilities are this model’s output, so the recipes stay in BF16, and the Playground offers FP8 for workloads that only use the top choice.
  • Reproducibility. Two launches of the same recipe returned bit-identical probabilities for the same requests at the same concurrency. Against the checkpoint’s reference implementation, the largest probability difference per question averages 0.003 (see §3), because batch composition and kernels differ, and near-ties at 0.500 can flip.

3. Benchmarks

All numbers in this section come from one NVIDIA GB300 (288 GB) at SGLang commit 70f0b7351e, with checkpoint revision 5117a6c, served with the GB300 recipe above unless a row says otherwise. Method. A closed-loop client keeps a fixed number of /v1/systemone requests in flight. Every request carries a unique random state, so requests share nothing but the system message, and each concurrency level starts after 8 warmup requests and a cache flush. Latency is the median end-to-end time of a request, and prompt tokens come from each response’s usage.input_tokens. Four request shapes cover the common cases: one four-option choice question about a 256-token state, four questions (two choice, one yes or no, one score) about a 1,024-token state, one question about an 8,192-token state, and one question about a 1024×768 JPEG with a 64-token state.

3.1 Speed

RequestConcurrentMedian latencyRequests/sPrompt tokens/s
Short question, 382 tokens151 ms19.47,413
4103 ms32.812,525
16267 ms59.022,531
64882 ms70.827,039
2563,521 ms71.127,164
Four questions on a 1K-token state, 4,540 tokens1179 ms5.6 (22 decisions/s)25,291
81,321 ms6.0 (24 decisions/s)27,332
325,206 ms6.0 (24 decisions/s)27,208
8K-token state, 8,224 tokens1325 ms3.0825,297
41,220 ms3.2626,785
164,796 ms3.3027,175
One 1024×768 image, 429 tokens187 ms11.54,945
16414 ms37.516,078
641,170 ms52.622,572
One GB300 tops out at about 27,000 prompt tokens per second for every text shape, so requests per second follow from prompt length. Latency rises with concurrency once that ceiling is reached, so choose the concurrency from the latency you can accept. A decision has no output tokens, so in the benchmark card above TTFT is the whole request latency, and TPOT and interactivity do not apply.

3.2 Accuracy and parity

Both sets go through /v1/systemone as one choice question per item. Belebele sends the passage as the state, the question as the instructions, and the four answers as options 1 to 4. WinoGrande sends the sentence as the state, asks which option fills the blank, and offers the two candidates. The reference row runs the same items through the checkpoint’s own DecisionModel one at a time.
ImplementationBelebele (eng_Latn, 900)WinoGrande (xl dev, 1,267)
SGLang, GB300 recipe96.67%84.21%
SGLang, no extra flag96.67%84.37%
Checkpoint’s reference DecisionModel (Transformers, BF16, GB300)96.67%84.37%
Perplexity, model card (Perplexity API, its own prompt conversion)94.00%83.30%
Against the reference implementation, the recipe picks the same option on 2,161 of 2,167 items, and each of the six differences is a near-tie that one side scores 0.500 against 0.500 and the other 0.486 against 0.514. Per item, the largest probability difference averages 0.003 (99th percentile 0.024, maximum 0.054). Perplexity’s figures come from a different prompt conversion, so they are context rather than a target.

3.3 What else was measured

Change from the GB300 recipeMeasured effectVerdict
Prefix cache on (no extra flag)Short question 58 ms alone, 8K-token state 377 ms alone and 2.87 requests/s saturated (recipe: 51 ms, 325 ms, 3.30)Only for traffic that repeats identical requests
—linear-attn-prefill-backend flashinferText 4-8% faster (8K-token state 300 ms, 3.56 requests/s), images slower (101 ms alone, 9% less throughput)Text-only traffic
—dp-size 2 (2 GPUs)139 short questions/s and 6.5 8K-token states/s, four questions on a 1K-token state in 104 ms, a lone short question 15 ms slowerScale-out
—tp 2 (2 GPUs)8K-token state 208 ms alone, 124 short questions/s saturated, a lone short question in 68 msLong states with a latency target
—quantization fp8About 20% more throughput, same accuracy, probabilities shifted by 0.010 on average and up to 0.20Only when you use the top choice alone
—chunked-prefill-size 32768 —max-prefill-tokens 32768No gain (within 4%)Keep the default
—mamba-ssm-dtype bfloat16No gain, though it halves the state poolKeep the default
Example
Output
Run it as python bench.py <state words> <concurrency> <requests>. With 270 words, a request is 378 tokens, close to the short shape above.
Example
Output

4. Advanced Usage

4.1 Text Decisions

Example
Output

4.2 Image Decisions

The example output below came from a mostly red test image.
Example
Output