Skip to main content

1. Model Introduction

Wan-Animate-2 is a 14B character-animation model in the Wan2.2 family. Instead of the text- or image-only conditioning of the other Wan2.2 checkpoints, it animates a reference image so that it follows the motion of a reference video, producing a new clip of that subject performing the reference video’s motion. Use Wan-AI/Wan2.2-Animate-2-14B-Diffusers as --model-path. The distilled checkpoint (Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers, log_scale=-1.3, 10 steps, no CFG) is not supported in this release. Unlike the A14B T2V/I2V checkpoints, Wan-Animate-2 is a single-expert model: it has one guidance_scale (default 3.0), no low-noise expert (guidance_scale_2) and no boundary switching. Diffusers >= 0.40.0 serves Wan-Animate-2 as a modular pipeline (WanAnimate2ModularPipeline); SGLang implements the model natively and loads the same Wan-AI/Wan2.2-Animate-2-14B-Diffusers checkpoint. There is no diffusers backend fallback.
Wan-Animate-2 is served through the native WanAnimate2Pipeline, selected automatically from the checkpoint’s model_index.json. It takes two inputs, a reference image and a reference video, so it is invoked differently from the other Wan2.2 models. See section 4.

1.1 Inputs

Both inputs are the base image_path / video_path sampling fields shared with the other video models; Wan-Animate-2 adds no model-specific input field. Offline runs pass both through a --config file (or --image-path / --video-path), and the online API forwards video_path into the pipeline while the reference image rides input_reference. Local paths must be readable by the server process. HTTP(S) video URLs use SGLang’s shared media loader before decoding, without a disk cache.

1.2 Defaults

Output length. The output has as many frames as the (resampled) reference video, not clip_len. The video is generated clip by clip with clip_len frames per chunk; a larger clip_len improves temporal coherence at the cost of VRAM. clip_len=81 roughly doubles the attention tokens of the default and OOMs a single 80 GB GPU. num_frames (and the API’s seconds, which maps to it) does not control the length: it is ignored and the server logs that it was, so leave it unset. Output resolution. width×height (the API’s size) is a pixel-area budget, not the output dimensions. The reference image’s aspect ratio is kept, the frame is the largest canvas with at most width × height pixels whose sides are multiples of 16, and the reference image and the driving-video frames are letterboxed onto it; the output MP4 has the canvas size. At the default 640×800 (512,000 pixels) a 9:16 reference gives 528×944, a 1:1 reference gives 704×704, and only a 4:5 reference gives exactly 640×800. To obtain a specific output size, request a budget with the same aspect ratio as the reference image (for a 9:16 image, 576×1024 maps onto itself). This differs from the OpenAI Videos API and from the other SGLang diffusion pipelines, where size is the literal output resolution. Audio. The output MP4 carries the reference video’s audio track, as the official pipeline does. Set enable_audio=false on a request (a --config field offline, a form field or extra_body entry online) for a silent output. If the track cannot be decoded, the output is silent and the server logs one warning per process; a reference video without an audio track also gives a silent output. Lossless conditioning reuse. Reference-video keys are rotated once and cached per clip, then reused by all denoising steps and both CFG branches. This removes repeated RoPE computation without changing the attention mask, precision, or sampling schedule. The cache uses the existing key storage and is released after the clip; it is not an approximate timestep-skipping cache. Clips containing only padding that would be removed from the final video are not denoised. The padding and conditioning of every retained clip are unchanged. Text encoding uses the native TextEncodingStage and its conditioning cache and residency management. CFG-parallel dispatch shares the native CFG utilities while retaining the single-GPU arithmetic order. Each clip uses the shared denoising loop for step profiling, NVTX ranges, progress, and DiT release. Reference K/V preparation and clip-specific prediction remain model hooks; clip decoding still feeds the next clip’s conditioning. This does not enable cache-dit or breakable CUDA graphs for this model. Clip decoding uses the shared DecodingStage.decode_raw path for VAE compilation, tiling and OOM diagnostics. Raw decoded pixels are kept for the next clip: no extra conversion to and from [0, 1] is introduced. The adapter preserves the existing FP32 VAE execution and per-channel latent scaling. QKV projections, cross-attention, FFN and the output head are inherited from the native Wan implementation. Ulysses exchanges use the shared packed-QKV path; the model-specific reference attention mask and K/V lifecycle are retained. Text and image conditioning projections are reused within each clip and CFG branch, rather than recomputed at every timestep. These projected tensors are released after the clip and never reused across requests. Decoded output is copied to CPU and clip-local intermediates are released; only the decoded overlap frames needed by the next clip remain on the GPU. Geometry caches retain at most four RoPE tables and two attention masks, rather than accumulating GPU tensors for every request size. Evicted entries are recomputed without changing the results; frequently alternating among more geometries can add setup work. Measured on one H200, with a 576×1024 budget, 8 steps, CFG 3, seed 42, clip_len=37, text-encoder CPU offload, and MP4 plus raw-frame output, the initial reuse changes reduced warm request latency from 42.30s to 41.34s (two requests per variant). With clip_len=17, omitting a fully cropped trailing clip reduced 59.60s to 40.45s (one warm request per variant). Raw RGB frames were bit-identical and peak allocated VRAM did not increase. These are workload-specific results, not general speedup claims; both comparisons used PyTorch 2.13.0+cu130, Transformers 5.12.1, and Diffusers 0.37.0. The baseline was the initial native implementation, not the official Diffusers implementation.

2. SGLang-diffusion Installation

SGLang-diffusion offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements. Please refer to the official SGLang-diffusion installation guide for installation instructions.

3. Model Deployment

3.1 Basic Configuration

On a single high-VRAM GPU, launch the native pipeline with:
Command
Routing to the native Wan-Animate-2 pipeline is automatic from the checkpoint’s model_index.json (_class_name: WanAnimate2Pipeline); pass --pipeline-class-name WanAnimate2Pipeline only for a local mirror that lacks model_index.json.

3.2 Configuration Tips

See Performance Optimization for acceleration features and their runtime requirements.
  • --num-gpus {NUM_GPUS}: Number of GPUs to use.
  • --tp-size {TP_SIZE}: Tensor parallelism size.
  • --ulysses-degree {ULYSSES_DEGREE}: DeepSpeed-Ulysses-style sequence parallelism (USP). The ring degree must be 1 for this model; --ring-degree greater than 1 is rejected with NotImplementedError. Combines with --tp-size: Ulysses splits the TP-local attention heads, so --tp-size must divide the 40 heads and --ulysses-degree must divide the per-shard count (see section 3.3).
  • --enable-cfg-parallel: Split the guided and unguided branches across GPUs.
  • --text-encoder-cpu-offload: CPU-offload the text encoder to save memory; it is moved to the GPU for the one encode per request.

3.3 Multi-GPU and memory presets

4. Model Invocation

4.1 Generate offline with a config file

The offline path passes both inputs (and the sampling parameters) through a --config YAML:
wan_animate_2_run.yaml
Command
Any value in the YAML can also be passed as a CLI flag (for example --image-path, --video-path, --num-inference-steps, --guidance-scale, --seed).

4.2 Serve and request over HTTP

POST /v1/videos creates an asynchronous job. Send the reference image as the input_reference part and the reference video as video_path, then poll the job and download the finished MP4. size is the pixel-area budget described in section 1.2; the output dimensions follow the reference image’s aspect ratio.
Command
The same request from the OpenAI Python client. The reference image is the input_reference file, and video_path (a path the server can read) and clip_len ride extra_body:
Python
For more API usage and request examples, see the SGLang Diffusion OpenAI API reference.

4.3 Reducing memory

When a single GPU is tight on VRAM, stream the DiT one layer at a time instead of resident-loading it:
Command
Combine with --text-encoder-cpu-offload --pin-cpu-memory to also keep the text encoder off the GPU between uses. Lowering clip_len also reduces peak memory.

5. Benchmark

Test Environment:
  • Hardware: NVIDIA GB200 GPU (1x, 2x and 4x; 189 GB per GPU)
  • Model: Wan-AI/Wan2.2-Animate-2-14B-Diffusers
  • sglang diffusion version: 0.5.22
  • Workload: size 640x800 (pixel-area budget) with a 16:9 reference image, so the frame is 960x528; a 237-frame reference video at fps 30 (7 clips of clip_len 37); 40 steps, CFG 3.0, seed 123; --text-encoder-cpu-offload --warmup-mode request. E2E is the request time after the warmup request, so it excludes the one-off flex-attention compile.
bench_serving cannot pass video_path, so the benchmark runs offline with sglang generate --config (the YAML from section 4.1) and reads timings from --perf-dump-path. SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 synchronizes the device so the per-stage timings are accurate.
Benchmark Command (1 GPU shown; add the flags from the table for the other rows):
Command
Result (peak memory per GPU from nvidia-smi):Layerwise DiT offload cuts the 1-GPU peak by 29 GiB at the same end-to-end time. The fastest 4-GPU layout is 3.5x faster than one GPU. The framework default and the explicit CFG-parallel x Ulysses layout run at the same speed and produce the same output; the explicit flags pin the layout. Tensor parallel x Ulysses is the memory-leaning 4-GPU layout: the lowest 4-GPU peak after --tp-size 4, about 13 percent slower than CFG-parallel x tensor parallel and 27 percent slower than the default. Adding --encoder-parallel replicate --vae-config.use-parallel-decode false to the CFG-parallel x Ulysses row costs about 1 percent (287 s, 62.5 GiB) and does not make the output bit-identical to a single GPU: every Ulysses layout matches the single-GPU video to about 30 dB mean PSNR. The layerwise rows are the configurations for 80 GB GPUs. With the replicated encoders, serial VAE decode and the channels-last VAE layout (SGLANG_DIFFUSION_VAE_CHANNELS_LAST_3D=1) the 2-GPU CFG-parallel output is bit-identical to a single GPU.

5.1 Native regression coverage

A separate regression used two B300 GPUs, CFG-parallel with replicated encoders and serial VAE decode, a 576x1024 area budget, 40 steps, clip_len=37, CFG=3, and seed 42. The driving input was looped to 64 frames to exercise two clips; this is a bounded regression workload, not a long-video quality benchmark. Both revisions (4ee69b4cf9ea and 04725c1c2900) used PyTorch 2.14.1+cu130, Transformers 5.19.0, Diffusers 0.37.0, and FlashAttention 4.0.0b34 on the same node. Cold and warm raw RGB outputs were byte-identical to the pre-optimization native implementation, both with resident DiT weights and with layerwise DiT offload. The single warm samples measured 112.29 s before and 112.40 s after optimization; offload measured 113.46 s. This does not establish a significant latency change. Rank 0’s request-stage peak allocated memory was 52.24, 52.23, and 23.93 GiB, respectively; these are not loading peaks or whole-device physical memory peaks. An H200 HTTP smoke test also covered image upload, asynchronous job polling and download, request warmup, repeated requests, URL video input, and reference-audio passthrough. Repeated decoded RGB outputs were identical. These checks preserve the existing native output, not bitwise parity with the newer Diffusers modular pipeline. The native sampler follows the original Wan sigma grid starting at exactly 1.0 and retains its BF16 conditioning conversion; the modular pipeline’s sampling and preprocessing differ. Official-reference accuracy must be validated separately with pinned versions and aligned settings.

6. Troubleshooting

  • OOM: lower clip_len (for example 37 instead of 65/81) or add --dit-layerwise-offload --layerwise-offload-components transformer.
  • --ring-degree greater than 1 is rejected at startup with NotImplementedError. Use --ulysses-degree for sequence parallelism and keep the ring degree at 1.
  • --ulysses-degree that does not divide the TP-local attention heads (40 / --tp-size; with TP 1, degrees 3, 6, 16, … are rejected) fails at startup with ValueError; with TP 1 use 2, 4, 5, 8, 10, 20 or 40. The same check covers leftover GPUs auto-assigned to sequence parallelism (for example --num-gpus 8 --tp-size 2 runs TP 2 x Ulysses 4).
  • FSDP (--use-fsdp-inference) is not supported for this model and is rejected at startup; use --tp-size or DiT layerwise offload for memory.
  • --attention-backend (or --component-attention-backends transformer=...) is rejected at startup with ValueError: the in-context self-attention runs on torch flex_attention only. The text encoder still accepts --component-attention-backends text_encoder=<backend>.
  • enable_teacache / enable_spectrum request fields are rejected with ValueError: every block runs at every step, so the cache heuristics do not apply.
  • SGLANG_DIFFUSION_ENABLE_MXFP8_ATTENTION is rejected at startup with ValueError: this model does not apply the offline Q/K rotation.
  • Silent output: the reference video’s audio track is kept only when it can be decoded; otherwise the output is silent and the server logs one warning per process naming the cause. A reference video without an audio track also gives a silent output (one info line). Pass enable_audio=false to skip audio extraction altogether.
  • Reference video not found: video_path must be a path the server process can read. Pass it as --video-path or through --config offline, or in the request body online.