> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sglang.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Hy3

## Introduction

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer
parameters, developed by the Tencent Hy Team. It features 192 experts with top-8 activation per token, GQA attention
(64 heads, 8 KV heads, head dim 128), and a 256K context length. The model supports built-in reasoning via
`reasoning_effort`, tool calling, and multi-token prediction (MTP) for speculative decoding. Hy3 significantly
outperforms similar-size models and rivals flagship open-source models with 2–5× the parameters, with notable
improvements in agentic workflows, long-context tasks, and production reliability.

This document demonstrates the deployment of Hy3 on Ascend NPUs using SGLang, including single-node (Atlas 800I A3) and
multi-node (Atlas 800I A2) PD mixed mode and speculative decoding.

## Supported features

| Feature              | Example usage                                                                                                                                                                                                              |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Tensor Parallelism   | `--tp-size 16`                                                                                                                                                                                                             |
| NPU Graph            | enabled by default; disable with `--disable-cuda-graph`;<br />control range via `--cuda-graph-bs`; e.g. `--cuda-graph-bs 4 8 16 20 24 28 32 36 40`                                                                         |
| Speculative Decoding | `--speculative-algorithm EAGLE \`<br />`--speculative-num-steps 2 \`<br />`--speculative-eagle-topk 1 \`<br />`--speculative-num-draft-tokens 3`                                                                           |
| Overlap Schedule     | `export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1`                                                                                                                                                                               |
| Reasoning Mode       | `extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}` — deep chain-of-thought;<br />`"reasoning_effort": "low"` — short thinking chain;<br />`"reasoning_effort": "no_think"` — direct response, no thinking |
| Tool Calling         | `--tool-call-parser auto --reasoning-parser auto`                                                                                                                                                                          |

<Note>
  The values in the **Example usage** column are for illustration only. Adjust them according to your hardware, deployment
  mode, and workload. For parameter details, see
  [Feature descriptions](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning#feature-descriptions).
</Note>

For feature compatibility and conflict information between features,
see [Feature Compatibility](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning#feature-compatibility).

## Prerequisites

### Environment

Before following this tutorial, complete the environment setup in the documents below:

* [Ascend NPU Quickstart](/docs/hardware-platforms/ascend-npus/getting-started/quick_start) — the fastest way to get started.
  It walks you through launching the official container image, starting the SGLang server, and sending a test request.
  Recommended if you are new to SGLang on Ascend.
* [SGLang Installation with NPU Support](/docs/hardware-platforms/ascend-npus/getting-started/installation) — the full installation guide.
  It covers the component version mapping (CANN, TorchNPU, Triton, kernels, etc.), building from source or from a
  Dockerfile, and recommended system settings (CPU power scheme, NUMA, swap). Use it when you need to install or customize
  the environment instead of using the official image.

### Model weights

<Warning>
  Hy3 has 295B total parameters. If you need to download model weights, check the model size before downloading to reserve enough space.
</Warning>

* [Hy3](https://www.modelscope.cn/models/Tencent-Hunyuan/Hy3) (BF16, 597.60GB)

<Tip>
  The BF16 variant (597.60GB) can be deployed on 16 × 64GB of device memory (`--tp-size 16`), which corresponds to one full
  Atlas 800I A3 node (8 cards × 2 dies) or two Atlas 800I A2 nodes (8 cards × 1 die each).
</Tip>

It is recommended to download the model weights to a shared directory accessible from within the container.

For the hardware specifications (memory per die, dies per card, and the difference between A2 and A3),
see [Ascend NPU Reference — Hardware](/docs/hardware-platforms/ascend-npus/reference/glossary#hardware).

## Installation

<Warning>
  Ensure sufficient disk space before pulling images. The Docker image requires at least **30GB** of free space.
</Warning>

The dependencies required for the NPU runtime environment have been integrated into a Docker image and uploaded to the
online platform. You can directly pull it.

<Note>
  Both **stable releases** and **daily builds** are available. The following command is based on the stable release tag.
  For details, see
  [Docker image versions](/docs/hardware-platforms/ascend-npus/faq#8-docker-image-versions-stable-release-vs-daily-build).
</Note>

<Tabs>
  <Tab title="Atlas 800I A3">
    ```bash Command theme={null}
    docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16

    docker run -itd --shm-size=64g --name ${NAME} \
    --privileged=true --net=host \
    -v /var/queue_schedule:/var/queue_schedule \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v /usr/local/sbin:/usr/local/sbin \
    -v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
    -v /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \
    --device=/dev/davinci0:/dev/davinci0  \
    --device=/dev/davinci1:/dev/davinci1  \
    --device=/dev/davinci2:/dev/davinci2  \
    --device=/dev/davinci3:/dev/davinci3  \
    --device=/dev/davinci4:/dev/davinci4  \
    --device=/dev/davinci5:/dev/davinci5  \
    --device=/dev/davinci6:/dev/davinci6  \
    --device=/dev/davinci7:/dev/davinci7  \
    --device=/dev/davinci8:/dev/davinci8  \
    --device=/dev/davinci9:/dev/davinci9  \
    --device=/dev/davinci10:/dev/davinci10  \
    --device=/dev/davinci11:/dev/davinci11  \
    --device=/dev/davinci12:/dev/davinci12  \
    --device=/dev/davinci13:/dev/davinci13  \
    --device=/dev/davinci14:/dev/davinci14  \
    --device=/dev/davinci15:/dev/davinci15  \
    --device=/dev/davinci_manager:/dev/davinci_manager \
    --device=/dev/hisi_hdc:/dev/hisi_hdc \
    --entrypoint=bash \
    quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
    ```
  </Tab>

  <Tab title="Atlas 800I A2">
    ```bash Command theme={null}
    docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16

    docker run -itd --shm-size=64g --name ${NAME} \
    --privileged=true --net=host \
    -v /var/queue_schedule:/var/queue_schedule \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v /usr/local/sbin:/usr/local/sbin \
    -v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
    -v /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \
    --device=/dev/davinci0:/dev/davinci0  \
    --device=/dev/davinci1:/dev/davinci1  \
    --device=/dev/davinci2:/dev/davinci2  \
    --device=/dev/davinci3:/dev/davinci3  \
    --device=/dev/davinci4:/dev/davinci4  \
    --device=/dev/davinci5:/dev/davinci5  \
    --device=/dev/davinci6:/dev/davinci6  \
    --device=/dev/davinci7:/dev/davinci7  \
    --device=/dev/davinci_manager:/dev/davinci_manager \
    --device=/dev/hisi_hdc:/dev/hisi_hdc \
    --entrypoint=bash \
    quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
    ```
  </Tab>
</Tabs>

<Tip>
  * If the model weights have already been downloaded to a shared directory, use `-v` to mount the model path into the
    container, for example: `-v /path/to/models:/models`.
  * Replace `${NAME}` with your own container name or remove `--name` to use default name.
  * `--shm-size=64g` is recommended for 16-NPU deployments with large batch sizes.
</Tip>

## Online service deployment

### Single-node online deployment

Single-node deployment completes both prefill and decode within the same node (PD mixed mode). Hy3 with 295B
parameters requires all 16 logical NPUs on an Atlas 800I A3 server (8 physical NPUs × 2 logical cores each, 64 GB HBM
per logical NPU).

Set the following environment variables before launching the server:

```bash theme={null}
# Performance tuning
export SGLANG_SET_CPU_AFFINITY=1
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0

# CANN environment (adjust paths to match your installation)
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh

# Ascend NPU settings
export ASCEND_USE_FIA=1
export STREAMS_PER_DEVICE=32
export HCCL_BUFFSIZE=3000
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo

# SGLang settings
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
export DEEP_NORMAL_MODE_USE_INT8_QUANT=1
```

Launch the server:

```bash theme={null}
MODEL_PATH="/path/to/Hy3"

python3 -m sglang.launch_server \
    --model-path ${MODEL_PATH} \
    --attention-backend ascend \
    --reasoning-parser auto \
    --tool-call-parser auto \
    --device npu \
    --tp-size 16 \
    --host 0.0.0.0 \
    --port 9999 \
    --mem-fraction-static 0.84 \
    --dtype bfloat16 \
    --base-gpu-id 0 \
    --prefill-max-requests 40 \
    --max-running-requests 40 \
    --cuda-graph-bs 4 8 16 20 24 28 32 36 40 \
    --speculative-algorithm EAGLE \
    --speculative-num-steps 2 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 3
```

### Multi-node PD mixed deployment on Atlas 800I A2

Each Atlas 800I A2 node provides 8 NPUs (Ascend 910B, 64 GB HBM per NPU). Hy3 requires all 16 logical NPUs across
**2 × Atlas 800I A2** nodes in PD mixed mode (prefill and decode are colocated on the same 2-node cluster).

On each node, set the following environment variables before launching the server:

```bash theme={null}
# ============================================================
# Before running, update the following variables:
#   HCCL_SOCKET_IFNAME: network interface name for HCCL (use `ifconfig` to find)
#   GLOO_SOCKET_IFNAME: network interface name for Gloo
# ============================================================

# Performance tuning
export SGLANG_SET_CPU_AFFINITY=1
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0

# CANN environment (adjust paths to match your installation)
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh

# Ascend NPU settings
export ASCEND_USE_FIA=1
export STREAMS_PER_DEVICE=32
export HCCL_BUFFSIZE=3000
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=<network-interface>
export GLOO_SOCKET_IFNAME=<network-interface>

# SGLang settings
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
export DEEP_NORMAL_MODE_USE_INT8_QUANT=1
```

Launch the server on both nodes by running the same script on each node. `NODE_IPS[0]` is the master node and must be
reachable from the other node:

```bash theme={null}
# ============================================================
# Before running, update the following variables:
#   MODEL_PATH: path to the model weights directory (shared)
#   NODE_IPS: IP addresses of the two nodes, e.g. ('10.0.0.1' '10.0.0.2')
# ============================================================

MODEL_PATH="/path/to/Hy3"
NODE_IPS=('<node1-ip>' '<node2-ip>')

for i in "${!NODE_IPS[@]}"; do
    if [[ "$(hostname -I | awk '{print $1}')" == "${NODE_IPS[$i]}" ]]; then
        python3 -m sglang.launch_server \
            --model-path ${MODEL_PATH} \
            --attention-backend ascend \
            --reasoning-parser auto \
            --tool-call-parser auto \
            --device npu \
            --tp-size 16 \
            --nnodes 2 \
            --dist-init-addr ${NODE_IPS[0]}:5000 \
            --node-rank $i \
            --host 0.0.0.0 \
            --port 9999 \
            --mem-fraction-static 0.84 \
            --dtype bfloat16 \
            --base-gpu-id 0 \
            --prefill-max-requests 40 \
            --max-running-requests 40 \
            --cuda-graph-bs 4 8 16 20 24 28 32 36 40 \
            --speculative-algorithm EAGLE \
            --speculative-num-steps 2 \
            --speculative-eagle-topk 1 \
            --speculative-num-draft-tokens 3
        break
    fi
done
```

<Tip>
  * `HCCL_SOCKET_IFNAME` and `GLOO_SOCKET_IFNAME` must be set to the cluster network interface on both nodes. Do **not** use
    `lo` for multi-node deployments.
  * Model weights must be accessible from both nodes, e.g. mounted from a shared directory.
  * Send requests only to the master node (`NODE_IPS[0]:9999`), not to the worker node.
</Tip>

## Functional verification

After the service is started, you can invoke the model by sending a prompt:

```shell theme={null}
# ============================================================
# Before running, update the following variables:
#   HOST: the server host address (e.g., localhost)
#   PORT: the server port number (e.g., 9999)
# ============================================================

curl http://${HOST}:${PORT}/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "Hy3",
        "messages": [{"role": "user", "content": "What is the capital of France?"}],
        "max_tokens": 64,
        "temperature": 0.9,
        "extra_body": {"chat_template_kwargs": {"reasoning_effort": "no_think"}}
    }'
```

Expected result: an HTTP 200 response with `"Paris"` in the content field.

To enable chain-of-thought reasoning for math or coding tasks, set `reasoning_effort` to `"high"`:

```shell theme={null}
curl http://${HOST}:${PORT}/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "Hy3",
        "messages": [{"role": "user", "content": "Solve: if 3x + 7 = 22, what is x?"}],
        "max_tokens": 8192,
        "temperature": 0.9,
        "extra_body": {"chat_template_kwargs": {"reasoning_effort": "high"}}
    }'
```

The response will contain a `reasoning_content` field with the thinking process and a `content` field with the final
answer.

Once the server prints `The server is fired up and ready to roll!` in the logs, it is ready to accept requests. For more testing examples (Health Check, Generate, Chat Completions, and port usage guidance), see [Testing the Service](/docs/hardware-platforms/ascend-npus/getting-started/installation#testing-the-service).

## Accuracy evaluation

For accuracy evaluation methods and datasets, see [Accuracy Evaluation on Ascend NPU](/docs/hardware-platforms/ascend-npus/evaluation/accuracy_evaluation).

## Performance

For performance data and benchmark commands, see [Performance Testing on Ascend NPU](/docs/hardware-platforms/ascend-npus/evaluation/performance_testing).

## Best practices

### Best practice configuration reference

Hy3 has no standalone best practice page yet. The [Single-node online deployment](#single-node-online-deployment) and
[Multi-node PD mixed deployment on Atlas 800I A2](#multi-node-pd-mixed-deployment-on-atlas-800i-a2) sections above provide
the ready-to-use scripts for the supported PD mixed cases.

For the full catalog of optimization features and their parameter and compatibility details,
see [Optimization on Ascend NPU](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning).

## Performance tuning

For the full list of supported features, see [Supported features](#supported-features). For detailed optimization
guidance, see [Optimization on Ascend NPU](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning).

### Reasoning mode selection

Hy3 supports three reasoning modes via `chat_template_kwargs`:

| `reasoning_effort` | Behavior                          | Recommended for                 |
| ------------------ | --------------------------------- | ------------------------------- |
| `no_think`         | No chain-of-thought, direct reply | Simple Q\&A, low-latency tasks  |
| `low`              | Short thinking chain              | Moderate reasoning tasks        |
| `high`             | Full chain-of-thought             | Math, coding, complex reasoning |

Use `no_think` for interactive or latency-sensitive applications to reduce output token count significantly.

## FAQ

For common environment, installation, and general parameter issues, please refer to the [Ascend NPU FAQ](/docs/hardware-platforms/ascend-npus/faq).
This section only covers model-specific issues.
