Skip to main content

Deployment

For all methods and hardware platforms, see the official SGLang installation guide. The two paths below match the Python / Docker toggle in the command panel. PPLX-Decider-v1.1 support lands in #42645, after v0.5.21, so until a release includes it, install a nightly build built after that merge.
Command
Then run the Python output of the command panel below in that environment.
Pick your GPU to generate the launch command. The model runs on one GPU in BF16 and needs no flag: SGLang reads the checkpoint’s decision_config.json, runs the full-attention layers over the whole prompt as the model was trained, and turns off the radix cache and chunked prefill, which would show those layers only part of the prompt.
The GB300 recipe was measured end to end on one GB300 at #42645: four request shapes from 1 to 256 concurrent requests, Belebele and WinoGrande accuracy, and per-question parity with the checkpoint’s own reference implementation. The H200 command is the same and has not been run yet. §3 has the method and the full tables.

Playground

The Playground layers SGLang features on top of the recipe above. Any override changes the badge to Not Verified until that exact configuration is tested end to end. The Configuration Tips say which overrides were measured on GB300. Tensor parallelism times data-parallel replicas must fit the GPUs on one node.

1. Model Introduction

PPLX-Decider-v1.1-27B is the successor to PPLX-Decider-v1-27B, Perplexity’s decision model fine-tuned from Qwen3.8-27B under the Apache-2.0 license. Like v1, it replaces the language model head with a readout over 255 answer codes and answers choice, yes or no (noul), and score questions about a text state and optional images with a probability for every option, without generating text. Two things changed. The 16 full-attention layers no longer use a causal mask, so every token attends to the whole prompt, while the 48 Gated DeltaNet layers stay causal. And training grew from 73,000 to 626,033 rows, most of them from tasksource. The prompt format, answer codes, and readout layout are the same as v1’s, but the calibrated temperature is 1.0087 instead of 2.2076, so probabilities from the two versions are not interchangeable.
PropertyValue
Checkpointperplexity-ai/pplx-decider-v1.1-27b (BF16)
Base modelQwen3.8-27B, one epoch of supervised fine-tuning on 626,033 decision rows (530,103 from tasksource)
AttentionNoncausal in the 16 full-attention layers, causal in the 48 Gated DeltaNet layers
Question typeschoice (up to 255 options), noul (yes or no), score (up to 10 levels)
InputsText or JSON state, optional images
Prompt lengthTrained on question prompts of up to 8,192 tokens, and the backbone accepts 262,144
LicenseApache-2.0
Perplexity reports a Decision Index of 61.56 for v1.1, up from 56.4 for v1:
Decision Index categoryJevPPLX-Decider-v1-27BPPLX-Decider-v1.1-27B
Knowledge51.440.948.18
Language62.063.569.45
Retrieval55.454.961.26
Tools75.179.378.88
Arts37.739.444.66
Overall57.956.461.56
Resources: HuggingFace, reference model code, Decision models in SGLang.

2. Configuration Tips

  • Route. Send requests to /v1/systemone, the System One API the checkpoint was trained for. /v1/decisions refuses this checkpoint because its labels and prompt are not the ones the readout learned, and chat or generation requests are not meaningful without a language model head.
  • Clients. Clients of the System One API, including the TypeSafe SDKs, work by pointing their base URL at the server. See System One compatible API for the request and response reference.
  • Images. Pass each image in images as base64 bytes, a base64 data URL, or an http(s) URL. They precede the text in every question. The processor resizes every image to between 65,536 and 262,144 pixels, so one image costs 64 to 256 prompt tokens.
  • Noncausal attention is automatic. SGLang reads attention_mode: noncausal_full_attention from decision_config.json and runs the full-attention layers as encoder layers, which the TRT-LLM, FlashInfer, and Triton attention backends support. It also turns off the radix cache and chunked prefill, logging Radix cache and chunked prefill are disabled for a decision checkpoint with noncausal full attention: a cached prefix or a prefill chunk would see only part of a prompt whose every position depends on the rest. These two settings override any flag you pass. A prompt longer than the 16,384-token prefill budget is still read in one pass: an 18,585-token prompt matched the reference implementation to four decimal places on GB300.
  • No prefix reuse. Because every layer output depends on the whole prompt, nothing carries over between requests, not even an identical one. The cost of a request is roughly questions × prompt tokens, as on v1.
  • Prompt length. Like v1, the model was trained on question prompts of up to 8,192 tokens, and the reference code refuses longer ones. SGLang serves longer prompts in one pass, but their accuracy is untested.
  • Attention backend. On GB300, Auto resolves to the TRT-LLM kernel with 64-token pages. --attention-backend flashinfer and --attention-backend triton matched the reference implementation as closely, with the same top option on 99.8% of the accuracy items. The H200 path has not been run yet.
  • Speed. One GB300 reads about 27,000 prompt tokens per second once batches are full, the same as v1. Measured back to back on the same GPU against v1’s recipe, v1.1 was within 2% on short questions and four-question requests, and an 8K-token state saturated 6% lower (3.13 against 3.32 requests per second), because noncausal attention computes the full attention matrix in its 16 full-attention layers instead of half of it.
  • More GPUs. The backbone and prompt sizes are v1’s, so the v1 page’s scaling measurements are the reference: data-parallel replicas (--dp-size N) for throughput, tensor parallelism only for latency on long prompts. They were not repeated for v1.1.
  • Upgrading from v1. Re-check any thresholds you tuned on v1 probabilities: v1.1 has a different temperature and changed its top answer on 8% of the Belebele and WinoGrande items.

3. Benchmarks

All numbers in this section come from one NVIDIA GB300 (288 GB) at #42645, with checkpoint revision 3b45dea, served with the GB300 recipe above unless a row says otherwise. The method and request shapes are those of the v1 page, whose speed script works here unchanged.

3.1 Speed

RequestConcurrentMedian latencyRequests/sPrompt tokens/s
Short question, 382 tokens153 ms18.97,203
4106 ms37.614,362
16275 ms57.421,909
64931 ms67.525,770
2563,489 ms71.127,139
Four questions on a 1K-token state, 4,540 tokens1183 ms5.5 (22 decisions/s)24,782
81,315 ms6.0 (24 decisions/s)27,398
324,882 ms6.1 (24 decisions/s)27,649
8K-token state, 8,224 tokens1334 ms2.9924,620
41,280 ms3.1225,682
165,101 ms3.1325,726
One 1024×768 image, 429 tokens1118 ms8.33,541
16466 ms33.614,398
641,197 ms49.921,416

3.2 Accuracy and parity

Both sets go through /v1/systemone as one choice question per item, with the same conversion as the v1 page. The reference rows run the same items through the checkpoint’s own DecisionModel one at a time.
ImplementationBelebele (eng_Latn, 900)WinoGrande (xl dev, 1,267)
SGLang, GB300 recipe (TRT-LLM attention)97.67%92.42%
SGLang, —attention-backend flashinfer97.67%92.58%
SGLang, —attention-backend triton97.67%92.42%
Checkpoint’s reference DecisionModel (Transformers, BF16, GB300)97.67%92.50%
v1.1 served with causal attention, as SGLang did before #4264597.00%88.87%
PPLX-Decider-v1-27B on SGLang, for comparison96.67%84.21%
Against the reference implementation, the recipe picks the same option on 2,164 of 2,167 items, and each of the three differences is a near-tie within 0.031 of an even split. Per item, the largest probability difference averages 0.002 (99th percentile 0.030, maximum 0.089). Prompts of 11,099 and 18,585 tokens, an image question, and a score question also matched the reference, the largest difference being 0.015 on the score question. Served with causal attention instead, v1.1 changed its top answer on 101 of the 2,167 items and lost 3.6 points on WinoGrande.
Example
Output

4. Advanced Usage

4.1 Text Decisions

Example
Output

4.2 Image Decisions

The example output below came from a 512×512 solid red test image.
Example
Output