- Chat-template hints (
enable_thinking,thinking_budget,reasoning_effort): the model’s own chat template decides what to do with them. No token guarantee. - Strict thinking (
--enable-strict-thinking): a grammar-level token filter that enforces an exact per-request budget. Recommended when you need a guarantee. - Custom logit processor (
--enable-custom-logit-processor): a logit processor that forces the thinking phase to end at the budget. Use when the grammar backend cannot be used.
Chat-template hints
The OpenAI chat completions API passeschat_template_kwargs through to the model’s chat template:
Chat-template kwargs
enable_thinkingtoggles the thinking mode for hybrid models such as Qwen3. It is template-dependent.thinking_budgetonly does something if the model’s own chat template defines that variable (the official Qwen3 template does). If the template does not define it, the kwarg is silently ignored.
reasoning_effort request field is forwarded into the chat template the same way.
These are hints to the model, not token guarantees. The model can overshoot the budget.
Hard budget with strict thinking
Launch the server with--enable-strict-thinking and a reasoning parser:
Launch with strict thinking
xgrammar is the default grammar backend, so no --grammar-backend flag is needed. Startup fails if the configured backend cannot filter tokens.
Enable --reasoning-parser so responses separate reasoning_content from the final answer (see Reasoning Parser). Strict thinking derives the thinking-end token ids from the parser’s think_end_token through the tokenizer, so it is not tied to any hardcoded ids.
Behavior
- When the budget is spent during thinking, the vocab mask allows only the
</think>token sequence, so the cap is exact. The model then produces the final answer. - With strict thinking enabled and no budget set, behavior is unchanged except that model-specific excluded tokens (for Qwen3,
<tool_call>,</tool_call>,<|im_end|>,<|endoftext|>) are blocked during the thinking phase. - To apply a server-wide cap to every request, set the environment variable
SGLANG_MAX_THINK_TOKENS(default-1, no cap).
Native /generate API
Set max_thinking_tokens per request, and always pair it with require_reasoning: true:
/generate with a thinking budget
OpenAI chat completions API
There is nomax_thinking_tokens field on chat completions. Pass the budget through custom_params instead; the runtime picks up the thinking_budget key the same way. The chat path sets require_reasoning for you based on the thinking mode:
Chat completions with a thinking budget
Hard budget with a custom logit processor
When the grammar backend cannot be used, a custom logit processor can enforce the budget instead. Launch the server with--enable-custom-logit-processor:
Launch with custom logit processors
/generate and chat completions:
Logit processor with a thinking budget
<think> start token. Once the budget is reached, it first forces a newline token, then forces </think>. The model may emit its own </think> right after the forced one; this is harmless.
Built-in processors
SGLang ships processors inpython/sglang/srt/sampling/custom_logit_processor.py, each with hardcoded thinking start / end / newline token ids:
These only work for models whose tokenizer produces those exact ids. For any other model, subclass
ThinkingBudgetLogitProcessor with the model’s own ids:
Custom thinking-budget processor
Verify the budget applies
Budget knobs can be silently ignored, so confirm the cap with a deterministic comparison:- Run a prompt with
temperature=0and no budget. Record the thinking length. - Run the same prompt with the budget set, again at
temperature=0.
- Missing
require_reasoning: trueon the native/generateAPI. chat_template_kwargs.thinking_budgeton a model whose chat template does not define the variable.- A built-in logit processor whose hardcoded token ids do not match the model’s tokenizer.
