Deployment
Install SGLang
Install SGLang
For all install methods and hardware platforms, see the official SGLang installation guide.
- Python (pip / uv)
- Docker
GigaChat 3.5 support is on Then run the Python output of the command panel below.
main (PR #29189, merged 2026-09-21) and not yet in a pip release — install from source:Command
- Low-Latency — MTP speculative decoding on, Mamba state pool pinned to 80 requests. Fastest reply per user: measured 2× decode speed at batch 1 and 1.5× at 16 concurrent 8k-token requests. Pick for chat and agent loops.
- High-Throughput — MTP off, pools left to the default ratio, so the KV pool is ~2.7× larger. Past a few dozen concurrent long-context requests the pinned recipe starts queueing prompts on KV space, and at saturation the draft + verify overhead no longer pays; this is the recipe for batch jobs and many concurrent users.
H100 numbers were measured on 8×H100 SXM at
b63f8416b3 (the PR #29189 merge on main) with bench_serving, random
ISL 8192 / OSL 1024, --random-range-ratio 1.0, --warmup-requests 4, --flush-cache; latencies are
P50. GSM8K was measured with the 5-shot chat harness sglang.test.run_eval --eval-name gsm8k --api chat (1314 scored questions,
max_tokens 8192), so the model thinks before every answer exactly as in production. That harness was
deprecated upstream after these runs; the Reproduce button emits its successor, sgl-eval run gsm8k
(single-shot), which can land a point or two away from the numbers shown.Playground
The Playground is where you experiment with SGLang features beyond the verified matrix. The Deploy panel above only emits combinations that have been run end-to-end; the Playground lets you turn on additional knobs — the reasoning and tool-call parsers, the MTP preset, parallelism — on top of whichever cell the Deploy panel is currently showing. Any change flips the badge to Not Verified until the new configuration is run end-to-end.1. Model Introduction
GigaChat 3.5 Reasoning is the thinking release of GigaChat’s hybrid Mixture-of-Experts model — 432B total parameters, 28B active per token, 256K context (YaRN, factor 8 over a 32K base), released under the MIT license. It shares the base architecture with GigaChat 3.5: 40 layers interleaving 10 full-attention MLA layers (DeepSeek-style latent attention, every 4th layer) with 30 GDN linear-attention layers (Qwen3-Next gated delta-net, constant-size state), and an MoE block routing to 8 of 256 experts plus 1 shared expert. Only the MLA layers hold a KV cache; the linear layers keep their state in a Mamba-style pool. What differs: the chat template opens every assistant turn with<think>, so the model always reasons before answering (there is no template switch to turn it off), and the checkpoint ships 3 stacked MTP draft heads instead of 2. Tool calls use the same GCML format. The model card reports the largest gains over the instant release in mathematics, code, instruction following and structured output.
Recommended generation: the model card suggests temperature=0.6 with a generous max_tokens (thinking alone can run to thousands of tokens). Informational only — the sample code below relies on the checkpoint’s generation_config.json defaults.
Resources: HuggingFace — GigaChat3.5-432B-A28B-Reasoning (FP8, served here) · GigaChat3.5-432B-A28B-Reasoning-bf16 (BF16, for fine-tuning; no single-node recipe).
2. Configuration Tips
Two pools. GDN layers keep per-request state in a Mamba-style pool, MLA layers use a paged KV pool;--mamba-full-memory-ratio (default 0.9) splits the post-weight budget. On 8×H100 at --mem-fraction-static 0.8 that is ~51 GB of weights per GPU, 348 state slots and 542k KV tokens. The resolver forces --mamba-radix-cache-strategy no_buffer (3 slots per request, overlap scheduler off), so the unpinned server admits 116 concurrent requests.
MTP needs a pinned pool and 0.83. Each running request holds extra state copies under speculative decoding, and the default ratio then admits only a few dozen. The Low-Latency recipe pins --max-mamba-cache-size 240 with --max-running-requests 80 and runs at --mem-fraction-static 0.83 (KV pool 199k tokens); at 0.85 the server ran out of memory in eager prefill at 16 concurrent 8k-token prompts. To size by workload instead of a pin: r = (3 + draft_tokens) × 16 MB / (11.5 KB × L) for an average request length L, and with MTP keep running 8k-token requests under ~20, since the EAGLE worker holds each prompt’s hidden states outside both pools.
Capacity levers (toggles in the Playground). --mamba-ssm-dtype bfloat16 halves every state slot at no cost in accuracy or accept length — on the Low-Latency recipe the KV pool grows from 199k to 496k tokens. --kv-cache-dtype fp8_e4m3 doubles the KV pool but decodes ~2.3× slower on H100 and cannot be combined with MTP on 80 GB.
MTP flags. --speculative-algorithm EAGLE uses GigaChat’s own NextN heads through the multi-layer EAGLE worker, selected automatically. This checkpoint has 3 heads: --speculative-num-steps 3, --speculative-num-draft-tokens 4.
Parsers. --reasoning-parser gigachat35 splits thinking from the answer (§3.1), --tool-call-parser gigachat35 gives structured tool calls (§3.2); auto resolves to gigachat35 for both. Do not use the reasoning parser on the Instant checkpoint.
3. Advanced Usage
3.1 Reasoning
The chat template opens every assistant turn with<think>, and there is no template switch to turn thinking off. Enable the gigachat35 reasoning parser (toggle Reasoning Parser in the Parsers card of the Playground above) to split the thinking from the final answer into reasoning_content vs content. Without the parser, the raw think block is returned inline in content.
Reasoning Example (Python)
Reasoning Example (Python)
Example
Example Output
Example Output
Output
The model can spend hundreds of tokens thinking, so keep
max_tokens generous. A response with an empty content and finish_reason: "length" means the think block was cut off before it closed — raise the limit.3.2 Tool Calling
Enable thegigachat35 tool-call parser (toggle Tool Call Parser in the Parsers card of the Playground above) to surface structured calls via message.tool_calls. GigaChat emits tool calls in its GCML format — a <|GCML|tool_calls> … </|GCML|tool_calls> block holding one <|GCML|invoke name="…"> per call — inside the assistant turn; the parser reads them out and strips the block from content. --tool-call-parser auto also resolves to gigachat35 for this checkpoint, since its chat template carries the GCML marker.
Tool Calling Example (Python)
Tool Calling Example (Python)
Example
Example Output
Example Output
Output
tool message with {"city": "Paris", "temp_c": 9, "condition": "light rain"} and re-requesting returns the final answer with no further tool call:Output
The tool-call turn carries the model’s thinking in
reasoning_content, so print it alongside content and tool_calls. GigaChat tends to call tools sequentially — one hop, then the next — so a multi-tool request may come back with a single call. Drive it with an agent loop that appends each tool result and re-requests until finish_reason is no longer tool_calls. Parallel calls (several invokes in one block) are supported by the format when the model chooses them.3.3 MTP (Speculative Decoding)
GigaChat 3.5 ships its own NextN draft heads, so speculative decoding needs no separate draft model:--speculative-algorithm EAGLE runs the stacked heads through SGLang’s multi-layer EAGLE worker, which is selected automatically for this architecture. The Low-Latency recipe has it on; in the Playground the Speculative card adds the same preset to any other cell.
Flags
- Steps follow the checkpoint. This checkpoint has 3 draft heads, so use
--speculative-num-steps 3and--speculative-num-draft-tokens 4(steps + 1); the worker rejects other values at startup. - Pin the state pool. Under MTP every running request holds extra GDN state copies for verification. Without
--max-mamba-cache-sizeand--max-running-requeststhe resolver clamps concurrency to a few dozen requests; §2 explains what the pin costs in context and why the recipe runs at--mem-fraction-static 0.83. The panel shows an amber callout whenever a speculative command lacks--max-running-requests. - Where it pays off. Long think blocks are long, low-entropy decode, exactly what draft heads predict well: on 8×H100 single-request decode runs about 2× faster and 16 concurrent 8k-token requests about 1.5× faster, with GSM8K unchanged. At 64 concurrent requests the verify work and the smaller KV pool make it slower than the plain recipe, which is why the High-Throughput recipe leaves MTP off. Measured numbers are in the benchmark card above.
