1. Model Introduction
Wan-Animate-2 is a 14B character-animation model in the Wan2.2 family. Instead of the text- or image-only conditioning of the other Wan2.2 checkpoints, it animates a reference image so that it follows the motion of a reference video, producing a new clip of that subject performing the reference video’s motion. UseWan-AI/Wan2.2-Animate-2-14B-Diffusers as --model-path. The distilled
checkpoint (Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers, log_scale=-1.3,
10 steps, no CFG) is not supported in this release.
Unlike the A14B T2V/I2V checkpoints, Wan-Animate-2 is a single-expert model:
it has one guidance_scale (default 3.0), no low-noise expert
(guidance_scale_2) and no boundary switching. Diffusers >= 0.40.0 serves
Wan-Animate-2 as a modular pipeline (WanAnimate2ModularPipeline); SGLang
implements the model natively and loads the same
Wan-AI/Wan2.2-Animate-2-14B-Diffusers checkpoint. There is no diffusers backend
fallback.
1.1 Inputs
Both inputs are the base
image_path / video_path sampling fields shared with the
other video models; Wan-Animate-2 adds no model-specific input field. Offline runs
pass both through a --config file (or --image-path / --video-path), and the
online API forwards video_path into the pipeline while the reference image rides
input_reference. Local paths must be readable by the server process. HTTP(S)
video URLs use SGLang’s shared media loader before decoding, without a disk cache.
1.2 Defaults
Output length. The output has as many frames as the (resampled) reference
video, not
clip_len. The video is generated clip by clip with clip_len frames
per chunk; a larger clip_len improves temporal coherence at the cost of VRAM.
clip_len=81 roughly doubles the attention tokens of the default and OOMs a
single 80 GB GPU. num_frames (and the API’s seconds, which maps to it) does
not control the length: it is ignored and the server logs that it was, so leave
it unset.
Output resolution. width×height (the API’s size) is a pixel-area budget,
not the output dimensions. The reference image’s aspect ratio is kept, the frame is
the largest canvas with at most width × height pixels whose sides are multiples of
16, and the reference image and the driving-video frames are letterboxed onto it;
the output MP4 has the canvas size. At the default 640×800 (512,000 pixels) a
9:16 reference gives 528×944, a 1:1 reference gives 704×704, and only a 4:5
reference gives exactly 640×800. To obtain a specific output size, request a
budget with the same aspect ratio as the reference image (for a 9:16 image,
576×1024 maps onto itself). This differs from the OpenAI Videos API and from the
other SGLang diffusion pipelines, where size is the literal output resolution.
Audio. The output MP4 carries the reference video’s audio track, as the
official pipeline does. Set enable_audio=false on a request (a --config field
offline, a form field or extra_body entry online) for a silent output. If the
track cannot be decoded, the output is silent and the server logs one warning per
process; a reference video without an audio track also gives a silent output.
Lossless conditioning reuse. Reference-video keys are rotated once and cached per clip, then reused by all
denoising steps and both CFG branches. This removes repeated RoPE computation
without changing the attention mask, precision, or sampling schedule. The cache
uses the existing key storage and is released after the clip; it is not an
approximate timestep-skipping cache.
Clips containing only padding that would be removed from the final video are not
denoised. The padding and conditioning of every retained clip are unchanged.
Text encoding uses the native TextEncodingStage and its conditioning cache and
residency management. CFG-parallel dispatch shares the native CFG utilities while
retaining the single-GPU arithmetic order.
Each clip uses the shared denoising loop for step profiling, NVTX ranges, progress,
and DiT release. Reference K/V preparation and clip-specific prediction remain
model hooks; clip decoding still feeds the next clip’s conditioning. This does
not enable cache-dit or breakable CUDA graphs for this model.
Clip decoding uses the shared DecodingStage.decode_raw path for VAE compilation,
tiling and OOM diagnostics. Raw decoded pixels are kept for the next clip: no
extra conversion to and from [0, 1] is introduced. The adapter preserves the
existing FP32 VAE execution and per-channel latent scaling.
QKV projections, cross-attention, FFN and the output head are inherited from the
native Wan implementation. Ulysses exchanges use the shared packed-QKV path;
the model-specific reference attention mask and K/V lifecycle are retained.
Text and image conditioning projections are reused within each clip and CFG
branch, rather than recomputed at every timestep. These projected tensors are
released after the clip and never reused across requests. Decoded output is
copied to CPU and clip-local intermediates are released; only the decoded
overlap frames needed by the next clip remain on the GPU.
Geometry caches retain at most four RoPE tables and two attention masks, rather
than accumulating GPU tensors for every request size. Evicted entries are
recomputed without changing the results; frequently alternating among more
geometries can add setup work.
Measured on one H200, with a 576×1024 budget, 8 steps, CFG 3, seed 42,
clip_len=37, text-encoder CPU offload, and MP4 plus raw-frame output, the initial
reuse changes reduced warm request latency from 42.30s to 41.34s (two requests
per variant). With clip_len=17, omitting a fully cropped trailing clip reduced
59.60s to 40.45s (one warm request per variant). Raw RGB frames were bit-identical
and peak allocated VRAM did not increase. These are workload-specific results,
not general speedup claims; both comparisons used PyTorch 2.13.0+cu130,
Transformers 5.12.1, and Diffusers 0.37.0. The baseline was the initial native
implementation, not the official Diffusers implementation.
2. SGLang-diffusion Installation
SGLang-diffusion offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements. Please refer to the official SGLang-diffusion installation guide for installation instructions.3. Model Deployment
3.1 Basic Configuration
On a single high-VRAM GPU, launch the native pipeline with:Command
model_index.json (_class_name: WanAnimate2Pipeline); pass
--pipeline-class-name WanAnimate2Pipeline only for a local mirror that lacks
model_index.json.
3.2 Configuration Tips
See Performance Optimization for acceleration features and their runtime requirements.--num-gpus {NUM_GPUS}: Number of GPUs to use.--tp-size {TP_SIZE}: Tensor parallelism size.--ulysses-degree {ULYSSES_DEGREE}: DeepSpeed-Ulysses-style sequence parallelism (USP). The ring degree must be 1 for this model;--ring-degreegreater than 1 is rejected withNotImplementedError. Combines with--tp-size: Ulysses splits the TP-local attention heads, so--tp-sizemust divide the 40 heads and--ulysses-degreemust divide the per-shard count (see section 3.3).--enable-cfg-parallel: Split the guided and unguided branches across GPUs.--text-encoder-cpu-offload: CPU-offload the text encoder to save memory; it is moved to the GPU for the one encode per request.
3.3 Multi-GPU and memory presets
4. Model Invocation
4.1 Generate offline with a config file
The offline path passes both inputs (and the sampling parameters) through a--config YAML:
wan_animate_2_run.yaml
Command
--image-path, --video-path, --num-inference-steps, --guidance-scale,
--seed).
4.2 Serve and request over HTTP
POST /v1/videos creates an asynchronous job. Send the reference image as the
input_reference part and the reference video as video_path, then poll
the job and download the finished MP4. size is the pixel-area budget described in
section 1.2; the output dimensions follow the reference image’s
aspect ratio.
Command
input_reference file, and video_path (a path the server can read) and
clip_len ride extra_body:
Python
4.3 Reducing memory
When a single GPU is tight on VRAM, stream the DiT one layer at a time instead of resident-loading it:Command
--text-encoder-cpu-offload --pin-cpu-memory to also keep the text
encoder off the GPU between uses. Lowering clip_len also reduces peak memory.
5. Benchmark
Test Environment:- Hardware: NVIDIA GB200 GPU (1x, 2x and 4x; 189 GB per GPU)
- Model: Wan-AI/Wan2.2-Animate-2-14B-Diffusers
- sglang diffusion version: 0.5.22
- Workload:
size640x800 (pixel-area budget) with a 16:9 reference image, so the frame is 960x528; a 237-frame reference video atfps30 (7 clips ofclip_len37); 40 steps, CFG 3.0, seed 123;--text-encoder-cpu-offload --warmup-mode request. E2E is the request time after the warmup request, so it excludes the one-off flex-attention compile.
bench_serving cannot pass video_path, so the benchmark runs offline
with sglang generate --config (the YAML from section 4.1)
and reads timings from --perf-dump-path. SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1
synchronizes the device so the per-stage timings are accurate.
- NVIDIA GB200
Benchmark Command (1 GPU shown; add the flags from the table for the other rows):Result (peak memory per GPU from
Command
nvidia-smi):Layerwise DiT offload cuts the 1-GPU peak by 29 GiB at the same end-to-end time.
The fastest 4-GPU layout is 3.5x faster than one GPU. The framework default and the explicit CFG-parallel x Ulysses
layout run at the same speed and produce the same output; the explicit flags pin the layout.
Tensor parallel x Ulysses is the memory-leaning 4-GPU layout: the lowest 4-GPU peak after
--tp-size 4, about 13 percent slower than CFG-parallel x tensor parallel and 27 percent slower than the default.
Adding --encoder-parallel replicate --vae-config.use-parallel-decode false to the CFG-parallel x Ulysses row costs about 1 percent (287 s, 62.5 GiB) and does not make the output bit-identical to a single GPU: every Ulysses layout matches the single-GPU video to about 30 dB mean PSNR.
The layerwise rows are the configurations for 80 GB GPUs. With the replicated encoders, serial VAE decode and the
channels-last VAE layout (SGLANG_DIFFUSION_VAE_CHANNELS_LAST_3D=1) the 2-GPU CFG-parallel output is bit-identical to a single GPU.5.1 Native regression coverage
A separate regression used two B300 GPUs, CFG-parallel with replicated encoders and serial VAE decode, a 576x1024 area budget, 40 steps,clip_len=37, CFG=3,
and seed 42. The driving input was looped to 64 frames to exercise two clips;
this is a bounded regression workload, not a long-video quality benchmark.
Both revisions (4ee69b4cf9ea and 04725c1c2900) used PyTorch 2.14.1+cu130,
Transformers 5.19.0, Diffusers 0.37.0, and FlashAttention 4.0.0b34 on the same node.
Cold and warm raw RGB outputs were byte-identical to the pre-optimization native
implementation, both with resident DiT weights and with layerwise DiT offload.
The single warm samples measured 112.29 s before and 112.40 s after optimization;
offload measured 113.46 s. This does not establish a significant latency change.
Rank 0’s request-stage peak allocated memory was 52.24, 52.23, and 23.93 GiB,
respectively; these are not loading peaks or whole-device physical memory peaks.
An H200 HTTP smoke test also covered image upload, asynchronous job polling and
download, request warmup, repeated requests, URL video input, and reference-audio
passthrough. Repeated decoded RGB outputs were identical.
These checks preserve the existing native output, not bitwise parity with the
newer Diffusers modular pipeline. The native sampler follows the original Wan
sigma grid starting at exactly 1.0 and retains its BF16 conditioning conversion;
the modular pipeline’s sampling and preprocessing differ. Official-reference
accuracy must be validated separately with pinned versions and aligned settings.
6. Troubleshooting
- OOM: lower
clip_len(for example37instead of65/81) or add--dit-layerwise-offload --layerwise-offload-components transformer. --ring-degreegreater than 1 is rejected at startup withNotImplementedError. Use--ulysses-degreefor sequence parallelism and keep the ring degree at 1.--ulysses-degreethat does not divide the TP-local attention heads (40 /--tp-size; with TP 1, degrees 3, 6, 16, … are rejected) fails at startup withValueError; with TP 1 use 2, 4, 5, 8, 10, 20 or 40. The same check covers leftover GPUs auto-assigned to sequence parallelism (for example--num-gpus 8 --tp-size 2runs TP 2 x Ulysses 4).- FSDP (
--use-fsdp-inference) is not supported for this model and is rejected at startup; use--tp-sizeor DiT layerwise offload for memory. --attention-backend(or--component-attention-backends transformer=...) is rejected at startup withValueError: the in-context self-attention runs on torchflex_attentiononly. The text encoder still accepts--component-attention-backends text_encoder=<backend>.enable_teacache/enable_spectrumrequest fields are rejected withValueError: every block runs at every step, so the cache heuristics do not apply.SGLANG_DIFFUSION_ENABLE_MXFP8_ATTENTIONis rejected at startup withValueError: this model does not apply the offline Q/K rotation.- Silent output: the reference video’s audio track is kept only when it can be decoded; otherwise the output is silent and the server logs one warning per process naming the cause. A reference video without an audio track also gives a silent output (one info line). Pass
enable_audio=falseto skip audio extraction altogether. - Reference video not found:
video_pathmust be a path the server process can read. Pass it as--video-pathor through--configoffline, or in the request body online.
