Skip to main content

1. Model Introduction

Hunyuan 3 Preview (Hy3-preview) is Tencent’s preview of its third-generation flagship MoE language model, featuring hybrid thinking, native tool calling, long-context reasoning, and Multi-Token Prediction (MTP) for low-latency serving. Key Features:
  • MoE Architecture: 192 routed experts + 1 shared expert, 8 experts activated per token. ~276B total parameters with ~20B active, delivering dense-model quality at MoE inference cost.
  • Hybrid Thinking: Reasoning modes (high, medium, low, none) controllable via OpenAI-standard reasoning_effort, allowing the same weights to trade off latency and depth of reasoning.
  • Native Tool Calling: Trained on structured <tool_call> / <arg_key> / <arg_value> grammar. Pairs with SGLang’s hunyuan tool-call parser for streaming OpenAI-compatible function-calling output.
  • Long Context: 256K token context window (262,144 positions) for repository-scale code and document reasoning.
  • Multi-Token Prediction (MTP): Ships with a built-in MTP draft module enabling speculative decoding out of the box.
Available Models: Recommended Generation Parameters:
ParameterValue
temperature0.7
top_p0.9
reasoning_efforthigh / medium / low (thinking) or none (instant)
License: TODO — verify on HuggingFace model card.

2. SGLang Installation

SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements. Please refer to the official SGLang installation guide for installation instructions. Docker Images by Hardware Platform:
Hardware PlatformDocker Image
NVIDIA H200 / B200 / B300 / GB300lmsysorg/sglang:latest
lmsysorg/sglang:latest bundles the HYV3 model code, the hunyuan tool-call / reasoning parsers, and the MTP draft-module runtime. For SGLang CPU installation, please refer to the CPU version installation guide.

3. Model Deployment

This section provides deployment configurations optimized for different hardware platforms and use cases.

3.1 Basic Configuration

Interactive Command Generator: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform, quantization, and feature capabilities.

3.2 Configuration Tips

Key Parameters:
ParameterDescriptionRecommended Value
--tool-call-parserTool call parser for function-calling supporthunyuan
--reasoning-parserReasoning parser for hybrid thinking modeshunyuan
--trust-remote-codeRequired for Hunyuan model loadingAlways enabled
--mem-fraction-staticStatic memory fraction (KV + activations)0.9
--tpTensor parallelism size2 / 4 / 8 depending on hardware
--attention-backendAttention backend (Blackwell only)trtllm_mha
--speculative-algorithmSpeculative decoding via the bundled MTP draftEAGLE + --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
Hardware Requirements: NVIDIA BF16 (Hy3-preview, ~552GB weights)
  • H200 (141GB) / B200 (180GB): TP=8 (minimum for BF16 to fit single-node).
  • B300 (275GB) / GB300: TP=4.
  • A100 / H100 (80GB): not supported single-node — BF16 requires multi-node TP=16+ on 80GB-class GPUs.
Blackwell (B200 / B300 / GB300): Auto-selected attention backend can mis-route for HYV3 on Blackwell. Always pass --attention-backend trtllm_mha explicitly on Blackwell hardware (the config generator above enforces this). Multi-Token Prediction (MTP): The Hy3-preview release bundles an MTP draft module. SGLang runs it via its EAGLE speculative-decoding path — the draft module auto-loads from the same --model-path. Enable with the standard MTP flags:
Command
Toggle the “Speculative Decoding (MTP)” option in the generator above to add these flags automatically. Tune num-steps / num-draft-tokens based on acceptance rate in your workload. Xeon CPU service configuration: Please refer to the Notes part in the serving engine launching section in the SGLang CPU server document to better understand how to configure the arguments, especially for TP (tensor parallel) and NUMA binding settings.

4. Model Invocation

4.1 Basic Usage

For basic API usage and request examples, please refer to: Deployment Command (H200 × 8, BF16 default):
Command
Testing Deployment: After startup, you can test the SGLang OpenAI-compatible API with the following command:
Command
Simple Completion Example:
Example
Output Example:
Output
When reasoning_effort is not set, the server defaults to instant mode (no thinking, reasoning_content=None). To opt into thinking, pass reasoning_effort="high" / "medium" / "low" on the request — see the Hybrid Thinking section below.

4.2 Advanced Usage

4.2.1 Reasoning Parser (Hybrid Thinking)

Hy3-preview is a hybrid-thinking model. Control the thinking budget via the OpenAI-standard reasoning_effort:
  • high / medium / low — increasing amounts of chain-of-thought in reasoning_content
  • none — skip thinking entirely (instant responses, content-only)
Enable the reasoning parser during deployment so that the thinking section (<think>...</think>) is separated into reasoning_content:
Command
Thinking Mode — High Effort:
Example
Output Example:
Output
Instant Mode — No Thinking:
Example
Output Example:
Output

4.2.2 Tool Calling

Hy3-preview supports streaming OpenAI-compatible tool calls. Enable both parsers together — the reasoning parser strips thinking tokens before the tool-call parser runs:
Command
Non-Streaming Example:
Example
Output Example:
Output
Streaming Example (incremental argument deltas): Hy3-preview’s hunyuan tool-call parser emits tool names first, then argument JSON in incremental fragments — matching the OpenAI streaming contract:
Example
Output Example:
Output

5. Benchmark

5.1 Accuracy Benchmark

Test Environment:
  • Hardware: 8× NVIDIA H200 (141GB)
  • Docker Image: lmsysorg/sglang:hy3-preview
  • Model: tencent/Hy3-preview (BF16)
  • Tensor Parallelism: 8
  • SGLang version: latest main

5.1.1 GSM8K

  • Benchmark Method: 5-shot CoT on 200 questions, evaluated via SGLang native backend
  • Benchmark Command:
Command
  • Test Results:
Output

5.1.2 MMLU

  • Benchmark Method: 5-shot, all 57 subjects
  • Benchmark Command:
Command
  • Test Results:
Output

5.1.3 Tool-Call Accuracy (MiniMax-Provider-Verifier)

  • Benchmark Tool: MiniMax-Provider-Verifier
  • Metric: function-call schema validity, argument match, and end-to-end response correctness
  • Test Results:
Output

5.2 Speed Benchmark

5.2.1 Low Concurrency

  • Benchmark Command:
Command
  • Test Results:
Output

5.2.2 High Concurrency

  • Benchmark Command:
Command
  • Test Results:
Output