Skip to main content

Deployment

Clef support merged after v0.5.21 in #42721. Use the pinned nightly wheel or Docker image below. See the installation guide for other methods.
This command selects the wheel from nightly build 37ae292e6f, which includes Clef support. It targets Python 3.10 on Linux x86_64 with glibc 2.39 or newer.
Command
Run the command from the wizard’s Python tab in this environment.
Choose your GPU, then Clef or Clef Flash. The wizard updates the model in the launch command and sample request. Both models run on one GPU in BF16 with TP=1 using the default settings; the commands need only the model path, host, and port.
We tested both models on one B300 and one H200 inside lmsysorg/sglang:dev-clef, built from SGLang bd2d73daa5af, with PyTorch 2.14.1+cu130 and Transformers 5.19.0. Each configuration passed the wizard’s cURL request, the Python example below, and four requests covering short and long text, each with and without an image. Both long requests used all 16,384 prompt tokens and preserved every question.B200 passed the same four text/image cases with SGLang e122069670a7, PyTorch 2.14.1, and Transformers 5.17.0. These checks covered serving behavior; throughput and accuracy were not measured.

1. Model introduction

Cloudflare post-trained Clef from Qwen3.8-27B and Clef Flash from Qwen3.5-9B, then released both under the Apache-2.0 license. Each model keeps its backbone and vision encoder and adds a joint schema head. The head reads the backbone’s final hidden states, routes state evidence to each option, and lets questions attend to one another. It returns one logit per allowed option in a single prefill, with no text generation.
VariantCheckpointBackbone
Clef (27B)Cloudflare/clefQwen3.8-27B
Clef Flash (9B)Cloudflare/clef-flashQwen3.5-9B
Pass a text string, JSON object, or JSON array as the state, with optional images. /v1/systemone returns one answer per question: a choice over allowed options, a noul probability of true, or a score over ordered levels.

2. Configuration tips

  • Route. Send requests to /v1/systemone, or set your System One client’s base URL to the server. This includes TypeSafe SDK clients. See the API reference for request and response fields. Clef refuses /v1/decisions, chat, and generation requests.
  • One GPU, one prefill per request. Keep TP=1 and PP=1. All questions share one prompt, so the model reads the state once. SGLang detects joint_head_config.json, loads joint_head.safetensors, and enables embedding mode. It disables radix caching, chunked prefill, and CUDA graphs. No reasoning parser, tool-call parser, or chat-template override is needed.
  • Answers. probabilities are the softmax of the head’s logits for each question. confidence uses the System One adapter formulas. It can differ from the checkpoint’s reference systemone(), which reports the top option’s probability as confidence. Answers omit x_label_mass because the head scores only the allowed options.
  • Long states. Prompts, including images, are capped at 16,384 tokens, the reference encode_record function’s default max_length. The effective limit can be lower, depending on the server context, KV token pool, and --max-prefill-tokens. The server truncates the state to leave room for the questions and images. If those do not fit even without the state, it rejects the request.
  • Images. Send each image in images as a data URL, such as data:image/jpeg;base64,..., or an HTTP(S) URL. Keep the data-URL prefix for JPEGs: bare base64 starts with /, which the loader treats as a file path, causing the request to fail. Images precede the state.

3. Advanced usage

3.1 Text decisions

Set MODEL to the checkpoint you selected in the wizard. The request works with either variant.
Example
Captured with Clef Flash on one H200. Probabilities can vary with the model variant, hardware, and numerical backend.
Output