Skip to main content

Deployment

For all install methods and hardware platforms, see the official SGLang installation guide.
GigaChat 3.5 support is on main (PR #29189, merged 2026-09-21) and not yet in a pip release — install from source:
Command
Then run the Python output of the command panel below.
This page covers the instant checkpoint (direct answers, no thinking). For the thinking checkpoint, see GigaChat 3.5 Reasoning. Pick your hardware and recipe to generate the launch command. Two serving strategies cover the operating points:
  • Low-Latency — MTP speculative decoding on, Mamba state pool pinned. Fastest reply per user; pick for chat and agent loops.
  • High-Throughput — MTP off. At saturation the draft + verify overhead outweighs the speedup, so this is the recipe for batch jobs and many concurrent users.

Playground

The Playground is where you experiment with SGLang features beyond the verified matrix. The Deploy panel above only emits combinations that have been run end-to-end; the Playground lets you turn on additional knobs — the tool-call parser, the MTP preset, parallelism — on top of whichever cell the Deploy panel is currently showing. Any change flips the badge to Not Verified until the new configuration is run end-to-end.

1. Model Introduction

GigaChat 3.5 is a hybrid Mixture-of-Experts model from GigaChat — 432B total parameters, 28B active per token, 256K context (YaRN, factor 8 over a 32K base), released under the MIT license. Its 40 layers interleave two attention types: 10 full-attention MLA layers (DeepSeek-style latent attention, every 4th layer) and 30 GDN linear-attention layers (Qwen3-Next gated delta-net, constant-size state). Only the MLA layers hold a KV cache; the linear layers keep their state in a Mamba-style pool. The MoE block routes to 8 of 256 experts plus 1 shared expert. The checkpoint ships 2 stacked MTP draft heads for speculative decoding, and tool calls are emitted in GigaChat’s GCML format. This is the instant release: it answers directly, without a thinking block. The separately trained Reasoning release always thinks first and ships 3 MTP heads — it has its own page. Resources: HuggingFace — GigaChat3.5-432B-A28B (FP8, served here) · GigaChat3.5-432B-A28B-bf16 (BF16, for fine-tuning; no single-node recipe).

2. Configuration Tips

Two pools. GDN layers keep per-request state in a Mamba-style pool, MLA layers use a paged KV pool; --mamba-full-memory-ratio (default 0.9) splits the post-weight budget. On 8×H100 at --mem-fraction-static 0.8 that is ~51 GB of weights per GPU, 347 state slots and 542k KV tokens. The resolver forces --mamba-radix-cache-strategy no_buffer (3 slots per request, overlap scheduler off), so the unpinned server admits 115 concurrent requests. MTP needs a pinned pool. Each running request holds extra state copies under speculative decoding, and the default ratio then admits only a few dozen. The Low-Latency recipe pins --max-mamba-cache-size 240 with --max-running-requests 80; the KV pool drops to 185k tokens. Keep --mem-fraction-static 0.8 on 80 GB — it is already tight in eager prefill at 16 concurrent 8k-token prompts. Capacity levers (toggles in the Playground). --mamba-ssm-dtype bfloat16 halves every state slot at no cost in accuracy or speed. --kv-cache-dtype fp8_e4m3 doubles the KV pool but decodes ~2.3× slower on H100 and cannot be combined with MTP on 80 GB. MTP flags. --speculative-algorithm EAGLE uses GigaChat’s own NextN heads through the multi-layer EAGLE worker, selected automatically. This checkpoint has 2 heads: --speculative-num-steps 2, --speculative-num-draft-tokens 3. Parsers. --tool-call-parser gigachat35 for structured tool calls (§3.1). No reasoning parser, not even auto: the checkpoint never emits </think>, and the forced gigachat35 splitter would move the whole answer into reasoning_content.

3. Advanced Usage

3.1 Tool Calling

Enable the gigachat35 tool-call parser (toggle Tool Call Parser in the Parsers card of the Playground above) to surface structured calls via message.tool_calls. GigaChat emits tool calls in its GCML format — a <|GCML|tool_calls> … </|GCML|tool_calls> block holding one <|GCML|invoke name="…"> per call — inside the assistant turn; the parser reads them out and strips the block from content. --tool-call-parser auto also resolves to gigachat35 for this checkpoint, since its chat template carries the GCML marker.
Example
Output
Appending a tool message with {"city": "Paris", "temp_c": 9, "condition": "light rain"} and re-requesting returns the final answer with no further tool call:
Output
GigaChat tends to call tools sequentially — one hop, then the next — so a multi-tool request may come back with a single call. Drive it with an agent loop that appends each tool result and re-requests until finish_reason is no longer tool_calls. Parallel calls (several invokes in one block) are supported by the format when the model chooses them.

3.2 MTP (Speculative Decoding)

GigaChat 3.5 ships its own NextN draft heads, so speculative decoding needs no separate draft model: --speculative-algorithm EAGLE runs the stacked heads through SGLang’s multi-layer EAGLE worker, which is selected automatically for this architecture. The Low-Latency recipe has it on; in the Playground the Speculative card adds the same preset to any other cell.
Flags
  • Steps follow the checkpoint. This checkpoint has 2 draft heads, so use --speculative-num-steps 2 and --speculative-num-draft-tokens 3 (steps + 1); the worker rejects other values at startup.
  • Pin the state pool. Under MTP every running request holds extra GDN state copies for verification. Without --max-mamba-cache-size and --max-running-requests the resolver clamps concurrency to a few dozen requests; §2 explains what the pin costs in context. The panel shows an amber callout whenever a speculative command lacks --max-running-requests.
  • Where it pays off. Low concurrency: accepted draft tokens skip target forwards, so single-request decode runs about 2× faster with the output distribution unchanged. At high concurrency the verify work and the smaller KV pool eat the gain, which is why the High-Throughput recipe leaves MTP off. Measured numbers are in the benchmark card above.