Deployment
Install SGLang
Install SGLang
Use an SGLang build that includes GLM-5.3-Flash support (v0.5.20 or later).For the listed NVIDIA hardware, the deployment panel can render a complete
Command
docker run command. See Install SGLang with Docker for host setup.For MI355X, follow the AMD GPU installation guide and run the generated Python command on four GPUs. Select FP8 or MXFP4; SGLang downloads the model and tokenizer automatically.- Low Latency starts with MTP 5/1/6 speculative decoding and tensor parallelism to shorten interactive responses.
- High Throughput starts with speculative decoding off, which avoids draft-and-verify overhead under sustained batches.
--cuda-graph-backend-prefill breakable to the generated command. This requires a build that includes PR #38522; the option is disabled on MI355X.
Generated commands leave
--mamba-full-memory-ratio at its 0.9 default, which is a generic starting point rather than a workload-tuned split: too low starves the KDA state pool and clamps max_running_requests, too high over-provisions it and shrinks the KV pool. Use the repo-local compute-mamba-ratio skill to compute the balanced ratio — or the --max-mamba-cache-size pin to use instead — from your average request length and the two pool sizes printed in one boot log. See Size both memory pools for what each pool caps.Playground
Use the Playground for lower-level tuning such as attention parallelism, MoE communication, speculative decoding, and reasoning or tool parsers. It inherits every selection from the deployment panel and shows only the command-line diff.1. Model introduction
GLM-5.3-Flash is a natively multimodal Mixture-of-Experts model built around a hybrid attention architecture — 320B total parameters with 18B active. Its 45 text layers combine MLA attention, DSA sparse attention, and KDA linear attention, and a 24-layer vision encoder handles image and video input. The checkpoint uses 288 routed experts with 8 active experts per token and includes a native MTP draft layer for speculative decoding. See the GLM-5.3-Flash blog and the GLM-5 technical report for training details.| Attribute | Description |
|---|---|
| Architecture | MoE with hybrid attention (MLA, DSA, KDA), mHC, and MTP — 320B total / 18B active parameters |
| Precision | FP8 weights (zai-org/GLM-5.3-Flash); MXFP4 on MI355X (OneNexus/GLM-5.3-Flash-MXFP4); NVFP4 on Blackwell from RadixArk/GLM-5.3-Flash-NVFP4 and nvidia/GLM-5.3-Flash-NVFP4; FP8 KV cache by default on Blackwell and MI355X, BF16 KV cache on H100 and H200 |
| Context | 1M tokens |
| Inputs | Text, image, and video |
| Generation defaults | temperature=1.0, top_p=0.95, thinking enabled |
2. Configuration tips
Choose a strategy
Start with Low Latency for chat and agent workloads. It drafts from the checkpoint’s MTP head at a fixed depth (5 steps, top-k 1, 6 draft tokens) with natural acceptance. Measure High Throughput for heavily batched traffic where disabling speculative decoding can be more efficient. SGLang serves MTP through--speculative-algorithm EAGLE (upstream folds the older NEXTN spelling into EAGLE), so generated commands use that flag value.
Strategy labels describe the workload goal, not a hardware restriction. Both strategies stay available when you switch hardware; only the verification badge changes. On MI355X, the measured AgentX frontier used Low Latency (EAGLE 5/1/6); High Throughput is an unmeasured speculation-off starting point.
Change the speculative algorithm
The Speculative card in the Playground changes the algorithm without leaving the selected strategy:- EAGLE / MTP 5-1-6 is exactly what Low Latency serves, so a Low Latency base starts on this chip. Pick it from a High Throughput base to keep that recipe’s other settings and add the MTP head.
- Off (greedy) strips the whole
--speculative-*family, which is what High Throughput already starts from. - DFlash2 swaps the in-checkpoint MTP head for the trained block-diffusion draft in
incoai/GLM-5.3-Flash-DFlash2. The draft proposes a whole block per step and the target verifies it in one forward pass, so output quality stays the target’s. Its block size comes from the draft checkpoint, and the draft runs onfa4rather than the target’s DSA backends. The hidden-state capture it needs (PR #36708) shipped with the GLM-5.3-Flash support in v0.5.20, so the image pinned above is enough. The draft repository is access-gated: request access on its model page, then download it alongside the target before serving. This combination is not yet measured on the cookbook hardware, so treat it as a starting point.
Size both memory pools
GLM-5.3-Flash maintains a paged KV pool for attention and a separate KDA state pool. The KDA state pool can limit concurrency before the KV pool is full. If startup reducesmax_running_requests because of KDA state capacity, increase --mamba-full-memory-ratio or set --max-mamba-cache-size for the expected concurrency, then tune --max-running-requests to the workload.
Keep the prefix cache enabled for every strategy.
Keep the checkpoint’s KDA lower-bound setting unchanged. In particular, do not override linear_lower_bound through --json-model-override-args.
Keep the KV and DSA backends paired
On Blackwell, the recipes default to an FP8 KV cache with TRT-LLM DSA: on GB300 this pairing measured 2.9–5.7% higher throughput and about 1.8x the KV token capacity at identical pool bytes, with GSM8K accuracy within noise of BF16. BF16 KV with TileLang DSA remains selectable in the deployment panel and is the default on H100 and H200, where FP8 KV with TRT-LLM DSA is disabled. MI355X instead uses FP8 KV with TileLang DSA in the recorded ROCm recipe. Switch the dtype and both DSA backends together; TileLang DSA with FP8 KV is not a valid CUDA combination, and TRT-LLM DSA is not the MI355X backend.Extend the cache hierarchy
Keep HiCache off when GPU memory is sufficient. Select L1 + L2 to spill reusable cache entries into host memory. Select + L3 only after configuring Mooncake on every serving node; the generated command exposes the required configuration path. These options remain selectable but are marked Not Verified until the resulting command is validated on the chosen hardware. For agentic MXFP4 serving on MI355X, select Low Latency with L1 + L2. The host tier uses--hicache-size 180 (about 190 GB of host memory per rank once the KDA state and DSA indexer pools are included, 750 GB per node) and requires a build that includes PR #42178. With less free host memory, keep HiCache off; the server does not start when the host tier does not fit.
Multimodal memory
All strategies enable multimodal serving. Image and video input need a Transformers build that shipsGlm5NextProcessor (the current transformers==5.17.0 dependency includes it and requires tokenizers>=0.23.1). With older transformers==5.12.1, the server loads only the tokenizer and silently answers image requests from the text alone. On MI355X, text, image, and an image-bearing follow-up were qualified with Transformers commit e4052f55; video was not evaluated. The processor samples video at 2 FPS and caps video input at 240,000 visual tokens. Install torchcodec in the serving environment before sending video requests. For very long videos on 4x GB300, use encoder disaggregation to isolate the vision encoder’s memory spikes from language decoding.
The default multimodal feature transport is automatic, and on a single CUDA node auto resolves to CPU transport. The MI355X selection explicitly emits CPU transport, matching its image-serving check. CUDA IPC is opt-in: pass --mm-feature-transport cuda_ipc when lower transfer latency matters more than the GPU memory the IPC pool reserves. CUDA VMM transport applies only to multi-node GB200/GB300 systems on the MNNVL fabric, where auto selects it.
3. Advanced usage
3.1 Reasoning
Thinking is enabled by the checkpoint’s generation configuration, and generated commands enable--reasoning-parser auto (which resolves to glm45 for GLM-5.3-Flash) by default. The OpenAI-compatible API then places thinking in message.reasoning_content and the final answer in message.content. You can disable Reasoning Parser in the Playground when an integration needs the raw response format.
To disable thinking for a request, pass chat_template_kwargs: {"thinking": false} in the request body.
3.2 Tool calling
Generated commands enable--tool-call-parser auto (which resolves to glm47 for GLM-5.3-Flash) by default, so structured calls are returned in message.tool_calls. You can disable Tool Call Parser in the Playground when tool calling is not needed. On follow-up turns, read both reasoning_content and content because a thinking model can use either field around tool execution.
3.3 Multimodal serving
The base recipes accept image and video content through the OpenAI-compatible chat API. Keep the processor defaults unless you have measured a different sampling or resize policy. Inputs above the video token budget are clamped to the processor’s limit rather than rejected.3.4 Encoder disaggregation
Encoder disaggregation separates vision preprocessing from language inference. The verified topology uses one 4x GB300 node shared by an encoder-only TP4 process on port 30001 and a language-only TP4 process on port 30000. Start the encoder first.Encoder server (GB300)
Encoder server (GB300)
Command
Language server (GB300)
Language server (GB300)
Command
--mem-fraction-static 0.78 on the language process so the encoder retains room for vision workspaces.
Install torchcodec for video and see the encoder disaggregation guide for the generic architecture and operational model.
3.5 Prefill Context Parallelism
Use prefill context parallelism (CP) to distribute DSA prefill computation for long prompts, including multimodal inputs. This hybrid model uses interleave CP for DSA layers and head-wise tensor parallelism for KDA layers, sharing the same communication group. Layer boundaries all-gather the context shards before KDA and reduce-scatter its output back onto those shards. The mHC residual streams and their coefficients stay local, with each output’s reduction completed before the residual update.Example of launching with Prefill CP on Blackwell GPU
Example of launching with Prefill CP on Blackwell GPU
Command
3.6 PD disaggregation (preview)
PD splits prefill and decode into separate server groups behind a router. For this hybrid model, the transfer moves both the paged DSA KV and the KDA recurrent state.PD serving on 4x GB300 (single node)
PD serving on 4x GB300 (single node)
Prefill (GPU 0-1)
Decode (GPU 2-3)
Router
8998 after --prefill must equal the prefill server’s --disaggregation-bootstrap-port.- Give each role a distinct
--nccl-portwhen both share one node. - Single-node NIXL needs
UCX_NET_DEVICES=loandUCX_TLS=tcp,cuda_copy,cuda_ipc,self,smin both server environments. - The validated arm used the triton MoE runner; the flashinfer_trtllm runner from the deployment recipes is untested under PD.
- Speculative decoding does not start under PD at the current cut (draft-graph capture width assert on the prefill role; decode-side memory pressure at TP2). Run PD without speculative flags.
- Prefill and decode with different TP sizes transfer state through slice paths, but numeric correctness is unverified. Keep both roles at the same TP size.
