Deployment
Install SGLang
Install SGLang
Clef support merged after v0.5.21 in #42721. Use the pinned nightly wheel or Docker image below. See the installation guide for other methods.
- Python (pip / uv)
- Docker
This command selects the wheel from nightly build 37ae292e6f, which includes Clef support. It targets Python 3.10 on Linux x86_64 with glibc 2.39 or newer.Run the command from the wizard’s Python tab in this environment.
Command
We tested both models on one B300 and one H200 inside
lmsysorg/sglang:dev-clef, built from SGLang bd2d73daa5af, with PyTorch 2.14.1+cu130 and Transformers 5.19.0. Each configuration passed the wizard’s cURL request, the Python example below, and four requests covering short and long text, each with and without an image. Both long requests used all 16,384 prompt tokens and preserved every question.B200 passed the same four text/image cases with SGLang e122069670a7, PyTorch 2.14.1, and Transformers 5.17.0. These checks covered serving behavior; throughput and accuracy were not measured.1. Model introduction
Cloudflare post-trained Clef from Qwen3.8-27B and Clef Flash from Qwen3.5-9B, then released both under the Apache-2.0 license. Each model keeps its backbone and vision encoder and adds a joint schema head. The head reads the backbone’s final hidden states, routes state evidence to each option, and lets questions attend to one another. It returns one logit per allowed option in a single prefill, with no text generation.| Variant | Checkpoint | Backbone |
|---|---|---|
| Clef (27B) | Cloudflare/clef | Qwen3.8-27B |
| Clef Flash (9B) | Cloudflare/clef-flash | Qwen3.5-9B |
/v1/systemone returns one answer per question: a choice over allowed options, a noul probability of true, or a score over ordered levels.
2. Configuration tips
- Route. Send requests to
/v1/systemone, or set your System One client’s base URL to the server. This includes TypeSafe SDK clients. See the API reference for request and response fields. Clef refuses/v1/decisions, chat, and generation requests. - One GPU, one prefill per request. Keep TP=1 and PP=1. All questions share one prompt, so the model reads the state once. SGLang detects
joint_head_config.json, loadsjoint_head.safetensors, and enables embedding mode. It disables radix caching, chunked prefill, and CUDA graphs. No reasoning parser, tool-call parser, or chat-template override is needed. - Answers.
probabilitiesare the softmax of the head’s logits for each question.confidenceuses the System One adapter formulas. It can differ from the checkpoint’s referencesystemone(), which reports the top option’s probability as confidence. Answers omitx_label_massbecause the head scores only the allowed options. - Long states. Prompts, including images, are capped at 16,384 tokens, the reference
encode_recordfunction’s defaultmax_length. The effective limit can be lower, depending on the server context, KV token pool, and--max-prefill-tokens. The server truncates the state to leave room for the questions and images. If those do not fit even without the state, it rejects the request. - Images. Send each image in
imagesas a data URL, such asdata:image/jpeg;base64,..., or an HTTP(S) URL. Keep the data-URL prefix for JPEGs: bare base64 starts with/, which the loader treats as a file path, causing the request to fail. Images precede the state.
3. Advanced usage
3.1 Text decisions
SetMODEL to the checkpoint you selected in the wizard. The request works with either variant.
Text decision example (Python)
Text decision example (Python)
Example
Example output
Example output
Captured with Clef Flash on one H200. Probabilities can vary with the model variant, hardware, and numerical backend.
Output
