Deployment
Install SGLang
Install SGLang
For all methods and hardware platforms, see the official SGLang installation guide. The two paths below match the Python / Docker toggle in the command panel. PPLX-Decider-v1.1 support lands in #42645, after v0.5.21, so until a release includes it, install a nightly build built after that merge.Then run the Python output of the command panel below in that environment.
- Python (pip / uv)
- Docker
Command
decision_config.json, runs the full-attention layers over the whole prompt as the model was trained, and turns off the radix cache and chunked prefill, which would show those layers only part of the prompt.
The GB300 recipe was measured end to end on one GB300 at #42645: four request shapes from 1 to 256 concurrent requests, Belebele and WinoGrande accuracy, and per-question parity with the checkpoint’s own reference implementation. The H200 command is the same and has not been run yet. §3 has the method and the full tables.
Playground
The Playground layers SGLang features on top of the recipe above. Any override changes the badge to Not Verified until that exact configuration is tested end to end. The Configuration Tips say which overrides were measured on GB300. Tensor parallelism times data-parallel replicas must fit the GPUs on one node.1. Model Introduction
PPLX-Decider-v1.1-27B is the successor to PPLX-Decider-v1-27B, Perplexity’s decision model fine-tuned from Qwen3.8-27B under the Apache-2.0 license. Like v1, it replaces the language model head with a readout over 255 answer codes and answers choice, yes or no (noul), and score questions about a text state and optional images with a probability for every option, without generating text.
Two things changed. The 16 full-attention layers no longer use a causal mask, so every token attends to the whole prompt, while the 48 Gated DeltaNet layers stay causal. And training grew from 73,000 to 626,033 rows, most of them from tasksource. The prompt format, answer codes, and readout layout are the same as v1’s, but the calibrated temperature is 1.0087 instead of 2.2076, so probabilities from the two versions are not interchangeable.
| Property | Value |
|---|---|
| Checkpoint | perplexity-ai/pplx-decider-v1.1-27b (BF16) |
| Base model | Qwen3.8-27B, one epoch of supervised fine-tuning on 626,033 decision rows (530,103 from tasksource) |
| Attention | Noncausal in the 16 full-attention layers, causal in the 48 Gated DeltaNet layers |
| Question types | choice (up to 255 options), noul (yes or no), score (up to 10 levels) |
| Inputs | Text or JSON state, optional images |
| Prompt length | Trained on question prompts of up to 8,192 tokens, and the backbone accepts 262,144 |
| License | Apache-2.0 |
| Decision Index category | Jev | PPLX-Decider-v1-27B | PPLX-Decider-v1.1-27B |
|---|---|---|---|
| Knowledge | 51.4 | 40.9 | 48.18 |
| Language | 62.0 | 63.5 | 69.45 |
| Retrieval | 55.4 | 54.9 | 61.26 |
| Tools | 75.1 | 79.3 | 78.88 |
| Arts | 37.7 | 39.4 | 44.66 |
| Overall | 57.9 | 56.4 | 61.56 |
2. Configuration Tips
- Route. Send requests to
/v1/systemone, the System One API the checkpoint was trained for./v1/decisionsrefuses this checkpoint because its labels and prompt are not the ones the readout learned, and chat or generation requests are not meaningful without a language model head. - Clients. Clients of the System One API, including the TypeSafe SDKs, work by pointing their base URL at the server. See System One compatible API for the request and response reference.
- Images. Pass each image in
imagesas base64 bytes, a base64 data URL, or anhttp(s)URL. They precede the text in every question. The processor resizes every image to between 65,536 and 262,144 pixels, so one image costs 64 to 256 prompt tokens. - Noncausal attention is automatic. SGLang reads
attention_mode: noncausal_full_attentionfromdecision_config.jsonand runs the full-attention layers as encoder layers, which the TRT-LLM, FlashInfer, and Triton attention backends support. It also turns off the radix cache and chunked prefill, loggingRadix cache and chunked prefill are disabled for a decision checkpoint with noncausal full attention: a cached prefix or a prefill chunk would see only part of a prompt whose every position depends on the rest. These two settings override any flag you pass. A prompt longer than the 16,384-token prefill budget is still read in one pass: an 18,585-token prompt matched the reference implementation to four decimal places on GB300. - No prefix reuse. Because every layer output depends on the whole prompt, nothing carries over between requests, not even an identical one. The cost of a request is roughly questions × prompt tokens, as on v1.
- Prompt length. Like v1, the model was trained on question prompts of up to 8,192 tokens, and the reference code refuses longer ones. SGLang serves longer prompts in one pass, but their accuracy is untested.
- Attention backend. On GB300, Auto resolves to the TRT-LLM kernel with 64-token pages.
--attention-backend flashinferand--attention-backend tritonmatched the reference implementation as closely, with the same top option on 99.8% of the accuracy items. The H200 path has not been run yet. - Speed. One GB300 reads about 27,000 prompt tokens per second once batches are full, the same as v1. Measured back to back on the same GPU against v1’s recipe, v1.1 was within 2% on short questions and four-question requests, and an 8K-token state saturated 6% lower (3.13 against 3.32 requests per second), because noncausal attention computes the full attention matrix in its 16 full-attention layers instead of half of it.
- More GPUs. The backbone and prompt sizes are v1’s, so the v1 page’s scaling measurements are the reference: data-parallel replicas (
--dp-size N) for throughput, tensor parallelism only for latency on long prompts. They were not repeated for v1.1. - Upgrading from v1. Re-check any thresholds you tuned on v1 probabilities: v1.1 has a different temperature and changed its top answer on 8% of the Belebele and WinoGrande items.
3. Benchmarks
All numbers in this section come from one NVIDIA GB300 (288 GB) at #42645, with checkpoint revision3b45dea, served with the GB300 recipe above unless a row says otherwise. The method and request shapes are those of the v1 page, whose speed script works here unchanged.
3.1 Speed
| Request | Concurrent | Median latency | Requests/s | Prompt tokens/s |
|---|---|---|---|---|
| Short question, 382 tokens | 1 | 53 ms | 18.9 | 7,203 |
| 4 | 106 ms | 37.6 | 14,362 | |
| 16 | 275 ms | 57.4 | 21,909 | |
| 64 | 931 ms | 67.5 | 25,770 | |
| 256 | 3,489 ms | 71.1 | 27,139 | |
| Four questions on a 1K-token state, 4,540 tokens | 1 | 183 ms | 5.5 (22 decisions/s) | 24,782 |
| 8 | 1,315 ms | 6.0 (24 decisions/s) | 27,398 | |
| 32 | 4,882 ms | 6.1 (24 decisions/s) | 27,649 | |
| 8K-token state, 8,224 tokens | 1 | 334 ms | 2.99 | 24,620 |
| 4 | 1,280 ms | 3.12 | 25,682 | |
| 16 | 5,101 ms | 3.13 | 25,726 | |
| One 1024×768 image, 429 tokens | 1 | 118 ms | 8.3 | 3,541 |
| 16 | 466 ms | 33.6 | 14,398 | |
| 64 | 1,197 ms | 49.9 | 21,416 |
3.2 Accuracy and parity
Both sets go through/v1/systemone as one choice question per item, with the same conversion as the v1 page. The reference rows run the same items through the checkpoint’s own DecisionModel one at a time.
| Implementation | Belebele (eng_Latn, 900) | WinoGrande (xl dev, 1,267) |
|---|---|---|
| SGLang, GB300 recipe (TRT-LLM attention) | 97.67% | 92.42% |
SGLang, —attention-backend flashinfer | 97.67% | 92.58% |
SGLang, —attention-backend triton | 97.67% | 92.42% |
Checkpoint’s reference DecisionModel (Transformers, BF16, GB300) | 97.67% | 92.50% |
| v1.1 served with causal attention, as SGLang did before #42645 | 97.00% | 88.87% |
| PPLX-Decider-v1-27B on SGLang, for comparison | 96.67% | 84.21% |
Check accuracy on your deployment (Python)
Check accuracy on your deployment (Python)
Example
Example Output
Example Output
Output
4. Advanced Usage
4.1 Text Decisions
Text Decision Example (Python)
Text Decision Example (Python)
Example
Example Output
Example Output
Output
4.2 Image Decisions
The example output below came from a 512×512 solid red test image.Image Decision Example (Python)
Image Decision Example (Python)
Example
Example Output
Example Output
Output
