1. Model Introduction
NVIDIA Cosmos3 is an omnimodal world-model family spanning text/image/video generation, optional synchronized sound, and robot action prediction. Its main advantage is breadth: the same native SGLang pipeline can serve media-generation checkpoints and the DROID policy checkpoint without routing through an LLM sampler. Choose Nano for the broadest modality coverage and lower deployment cost, Super for the larger 64B image/video model, and a specialized checkpoint when only T2I or I2V is needed. Sound and action are checkpoint-specific heads, so they are not available from every Cosmos3 repository.
Sound and action generation require the corresponding checkpoint heads. The pipeline reads the transformer and scheduler configs at startup, so Edge and distilled checkpoints do not require architecture-specific server flags. Non-distilled checkpoints use the flow-native
FlowUniPCMultistepScheduler; distilled checkpoints use the fixed sigma schedule stored in the checkpoint.
The default flow_shift is 3.0 for T2I, 10.0 for non-Edge video and all action modes, and 3.0 for Edge video modes. Distilled checkpoints bake the schedule into their sigmas and do not use a request-level flow_shift.
2. Installation
Install SGLang with the diffusion dependencies:Command
Command
cosmos-guardrail downloads gated NVIDIA guardrail weights, so pass a Hugging Face token if your environment needs one. If the package is not installed, SGLang skips Cosmos3 guardrails and logs a warning. To disable Cosmos3 guardrails for local experiments, set SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 before starting the server.
There may be problems loading the Cosmos-1.0-Guardrail weights on Ascend NPU. If the _pickle.UnpicklingError error occurs during startup, you should change weight_only=True to weights_only=False parameter in cosmos_guardrail/cosmos_utils.py:
3. Serve Cosmos3
ServeCosmos3-Nano directly from the Hugging Face model ID:
Command
--performance-mode auto, Cosmos3 Nano keeps its DiT and VAE resident
when every selected GPU has at least 90 GiB available at startup. Other
Cosmos3 checkpoints use a 120 GiB threshold. Below the applicable threshold,
auto mode retains the conservative DiT component-offload policy. Cosmos3 runs
one DiT per pipeline, so component offload above the threshold only pays to
copy the weights out to host memory and back on every request. Serve
Cosmos3-Super across multiple GPUs as shown below so each rank holds a shard
of the weights.
For Cosmos3-Super, split the model across multiple GPUs:
Command
nvidia/Cosmos3-Super-Text2Image and nvidia/Cosmos3-Super-Image2Video checkpoint IDs.
Edge checkpoints
Cosmos3-Edge is a 4B dense model and can be served on one GPU:
Command
832x480 with guidance_scale=5.0; its default image configuration is 640x640 with guidance_scale=7.0. Supported sizes are 832x480, 480x832, 640x480, 480x640, 480x480, 640x640, 448x256, 256x448, and 256x256.
Serve the Edge DROID policy checkpoint with the same single-GPU configuration, replacing the model path with nvidia/Cosmos3-Edge-Policy-DROID.
Distilled checkpoints
The distilled Super checkpoints are 64B models. Use multiple GPUs unless the complete model and request workload fit on one GPU:Command
nvidia/Cosmos3-Super-Image2Video-4Step. SGLang detects both checkpoints from scheduler/scheduler_config.json, uses the checkpoint’s fixed four-step sigma schedule, and forces guidance_scale=1.0. Do not tune num_inference_steps or flow_shift for these checkpoints.
4. OpenAI-Compatible Requests
Text to image
Cosmos3 text-to-image uses/v1/images/generations. The default Cosmos3 image response is b64_json, matching vLLM-Omni’s examples.
Command
nvidia/Cosmos3-Super-Text2Image-4Step, omit the scheduler controls and use guidance_scale=1.0:
Command
Text to video with sound
Use/v1/videos to create an asynchronous job, then poll the job and download the completed MP4. Set generate_sound=true to generate and mux a stereo 48 kHz audio track; omit it for a silent video.
Command
Image to video
This mirrors the officialnvidia/Cosmos3-Nano Hugging Face image-to-video example:
Python
nvidia/Cosmos3-Super-Image2Video-4Step. The recommended request is 480p and does not specify scheduler controls:
Command
Video to video
Upload a source video withvideo_reference. Cosmos3 keeps latent frames [0, 1] by default and generates the remaining frames. Use condition_frame_indexes to select different latent frames, and condition_video_keep to take conditioning frames from the start or end of the source.
Command
Action generation
For DROID policy generation, start a single-GPU server with either the Nano or Edge policy checkpoint. Cosmos3 action generation does not currently support CFG or sequence parallelism.Command
nvidia/Cosmos3-Edge-Policy-DROID in the same command to serve the smaller 4B policy checkpoint.
policy and inverse_dynamics return actions, so their canonical API is the synchronous /v1/actions/generations endpoint. The following request predicts a 16-step action chunk from one observation image. action_horizon=16 maps to the model’s num_frames=17 convention.
Python
GET /v1/actions/metadata to inspect the action modes, default horizon, padded action dimension, and accepted observation modalities. Msgpack requests and the /v1/actions/realtime websocket use the same action envelope.
To batch policy observations inside one request, opt in with a bounded batch size:
Command
[B, H, W, C] uint8 array in input.input_reference, and either one prompt per image or one scalar prompt to broadcast across the batch. Batched prompts must currently tokenize to the same length because Cosmos3 GEN cross-attention does not mask padded text K/V. All items in one request share the domain, resolution, action horizon, and denoise settings. The standard action envelope returns one data[i] item per input, each with action shape [H, D]. For a compact msgpack response containing one [B, H, D] array, set runtime.response_format="raw" and read the top-level actions field.
For JSON, input_reference can be a list of base64 image payloads. For msgpack, it can be a packed uint8 numpy array directly:
JSON
--batching-max-size; this keeps one request from bypassing the server’s configured memory limit. Batching applies to action_mode="policy" only. A request seed controls the random stream for the whole batch, so a batched result is deterministic for that request but is not expected to be bit-exact with separately seeded B=1 requests.
inverse_dynamics also uses /v1/actions/generations; set action_mode="inverse_dynamics" and pass an observation video URL or server-local path as input.observation.video. Select the embodiment head with domain_name or domain_id; set raw_action_dim explicitly when it cannot be inferred from the domain name.
forward_dynamics is intentionally different: it consumes an action array and predicts video, so it remains on /v1/videos. Action-producing modes submitted to /v1/videos return HTTP 400 with the canonical action endpoint in the error message.
5. Cosmos3 Parameters
Cosmos3 supports the standard SGLang video and image fields such assize, num_frames, fps, num_inference_steps, guidance_scale, negative_prompt, and seed. For distilled checkpoints, SGLang replaces num_inference_steps with the checkpoint’s fixed four-step schedule and forces guidance_scale=1.0; negative-prompt CFG and request-level flow_shift do not apply.
Top-level Cosmos3 request fields:
max_sequence_length: maximum text token length used by the Cosmos3 tokenizer.flow_shift: per-request scheduler shift for non-distilled checkpoints. If omitted, SGLang uses--flow-shift, then the mode default (3.0for T2I,10.0for non-Edge video and all action modes, or3.0for Edge video).guidance_interval: optional[start, end]noise interval for CFG. Non-distilled T2I defaults to[400, 1000]; video modes guide at every step.
generate_sound: generate a sound track whose duration followsnum_frames / fps.sound_duration: explicit sound duration in seconds; takes precedence over the derived duration.condition_frame_indexes: V2V latent-frame indexes to keep from the source video; defaults to[0, 1].condition_video_keep: use thefirstorlastsource frames for V2V conditioning.action_mode:policy,forward_dynamics, orinverse_dynamics.domain_name/domain_id: select the action embodiment head.raw_action_dim: number of active action dimensions; inferred for known domain names.action: action array with shape[T, D], required byforward_dynamics.action_fps: action-token frame rate for temporal mRoPE; defaults to the video FPS.action_view_point: viewpoint used in the structured action caption.action_normalization: dataset normalization mode, such asquantile,meanstd, orminmax.
extra_body with the OpenAI Python SDK.
Raw JSON may keep them at the top level; multipart video requests should put
them in the extra_params JSON object. The legacy image extra_args container
remains accepted for compatibility, but new clients should use extra_body:
use_duration_template: whether to append SGLang’s generated duration suffix to video prompts.use_resolution_template: accepted for vLLM-Omni request compatibility.use_system_prompt: whether to add the Cosmos3 system prompt to the chat template.guardrailsoruse_guardrails: per-request guardrail toggle when the server started with guardrails enabled.
