Skip to main content

1. Model Introduction

NVIDIA Cosmos3 is an omnimodal world-model family spanning text/image/video generation, optional synchronized sound, and robot action prediction. Its main advantage is breadth: the same native SGLang pipeline can serve media-generation checkpoints and the DROID policy checkpoint without routing through an LLM sampler. Choose Nano for the broadest modality coverage and lower deployment cost, Super for the larger 64B image/video model, and a specialized checkpoint when only T2I or I2V is needed. Sound and action are checkpoint-specific heads, so they are not available from every Cosmos3 repository. Sound and action generation require the corresponding checkpoint heads. The pipeline reads the transformer and scheduler configs at startup, so Edge and distilled checkpoints do not require architecture-specific server flags. Non-distilled checkpoints use the flow-native FlowUniPCMultistepScheduler; distilled checkpoints use the fixed sigma schedule stored in the checkpoint. The default flow_shift is 3.0 for T2I, 10.0 for non-Edge video and all action modes, and 3.0 for Edge video modes. Distilled checkpoints bake the schedule into their sigmas and do not use a request-level flow_shift.

2. Installation

Install SGLang with the diffusion dependencies:
Command
Cosmos3 guardrails are enabled by default when the package is available:
Command
cosmos-guardrail downloads gated NVIDIA guardrail weights, so pass a Hugging Face token if your environment needs one. If the package is not installed, SGLang skips Cosmos3 guardrails and logs a warning. To disable Cosmos3 guardrails for local experiments, set SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 before starting the server. There may be problems loading the Cosmos-1.0-Guardrail weights on Ascend NPU. If the _pickle.UnpicklingError error occurs during startup, you should change weight_only=True to weights_only=False parameter in cosmos_guardrail/cosmos_utils.py:

3. Serve Cosmos3

Serve Cosmos3-Nano directly from the Hugging Face model ID:
Command
With --performance-mode auto, Cosmos3 Nano keeps its DiT and VAE resident when every selected GPU has at least 90 GiB available at startup. Other Cosmos3 checkpoints use a 120 GiB threshold. Below the applicable threshold, auto mode retains the conservative DiT component-offload policy. Cosmos3 runs one DiT per pipeline, so component offload above the threshold only pays to copy the weights out to host memory and back on every request. Serve Cosmos3-Super across multiple GPUs as shown below so each rank holds a shard of the weights. For Cosmos3-Super, split the model across multiple GPUs:
Command
The server also accepts the specialized nvidia/Cosmos3-Super-Text2Image and nvidia/Cosmos3-Super-Image2Video checkpoint IDs.

Edge checkpoints

Cosmos3-Edge is a 4B dense model and can be served on one GPU:
Command
Edge is trained for 256p and 480p generation. Its default video configuration is 832x480 with guidance_scale=5.0; its default image configuration is 640x640 with guidance_scale=7.0. Supported sizes are 832x480, 480x832, 640x480, 480x640, 480x480, 640x640, 448x256, 256x448, and 256x256. Serve the Edge DROID policy checkpoint with the same single-GPU configuration, replacing the model path with nvidia/Cosmos3-Edge-Policy-DROID.

Distilled checkpoints

The distilled Super checkpoints are 64B models. Use multiple GPUs unless the complete model and request workload fit on one GPU:
Command
For distilled I2V, replace the model path with nvidia/Cosmos3-Super-Image2Video-4Step. SGLang detects both checkpoints from scheduler/scheduler_config.json, uses the checkpoint’s fixed four-step sigma schedule, and forces guidance_scale=1.0. Do not tune num_inference_steps or flow_shift for these checkpoints.

4. OpenAI-Compatible Requests

Text to image

Cosmos3 text-to-image uses /v1/images/generations. The default Cosmos3 image response is b64_json, matching vLLM-Omni’s examples.
Command
With a server running nvidia/Cosmos3-Super-Text2Image-4Step, omit the scheduler controls and use guidance_scale=1.0:
Command

Text to video with sound

Use /v1/videos to create an asynchronous job, then poll the job and download the completed MP4. Set generate_sound=true to generate and mux a stereo 48 kHz audio track; omit it for a silent video.
Command

Image to video

This mirrors the official nvidia/Cosmos3-Nano Hugging Face image-to-video example:
Python
For the distilled I2V checkpoint, use the same API with a server running nvidia/Cosmos3-Super-Image2Video-4Step. The recommended request is 480p and does not specify scheduler controls:
Command
Poll and download this job with the same status and content endpoints used by the T2V example.

Video to video

Upload a source video with video_reference. Cosmos3 keeps latent frames [0, 1] by default and generates the remaining frames. Use condition_frame_indexes to select different latent frames, and condition_video_keep to take conditioning frames from the start or end of the source.
Command
Poll and download this job with the same status and content endpoints used by the T2V example.

Action generation

For DROID policy generation, start a single-GPU server with either the Nano or Edge policy checkpoint. Cosmos3 action generation does not currently support CFG or sequence parallelism.
Command
Use nvidia/Cosmos3-Edge-Policy-DROID in the same command to serve the smaller 4B policy checkpoint. policy and inverse_dynamics return actions, so their canonical API is the synchronous /v1/actions/generations endpoint. The following request predicts a 16-step action chunk from one observation image. action_horizon=16 maps to the model’s num_frames=17 convention.
Python
Use GET /v1/actions/metadata to inspect the action modes, default horizon, padded action dimension, and accepted observation modalities. Msgpack requests and the /v1/actions/realtime websocket use the same action envelope. To batch policy observations inside one request, opt in with a bounded batch size:
Command
Send one image per observation as a list or [B, H, W, C] uint8 array in input.input_reference, and either one prompt per image or one scalar prompt to broadcast across the batch. Batched prompts must currently tokenize to the same length because Cosmos3 GEN cross-attention does not mask padded text K/V. All items in one request share the domain, resolution, action horizon, and denoise settings. The standard action envelope returns one data[i] item per input, each with action shape [H, D]. For a compact msgpack response containing one [B, H, D] array, set runtime.response_format="raw" and read the top-level actions field. For JSON, input_reference can be a list of base64 image payloads. For msgpack, it can be a packed uint8 numpy array directly:
JSON
The batch size cannot exceed --batching-max-size; this keeps one request from bypassing the server’s configured memory limit. Batching applies to action_mode="policy" only. A request seed controls the random stream for the whole batch, so a batched result is deterministic for that request but is not expected to be bit-exact with separately seeded B=1 requests. inverse_dynamics also uses /v1/actions/generations; set action_mode="inverse_dynamics" and pass an observation video URL or server-local path as input.observation.video. Select the embodiment head with domain_name or domain_id; set raw_action_dim explicitly when it cannot be inferred from the domain name. forward_dynamics is intentionally different: it consumes an action array and predicts video, so it remains on /v1/videos. Action-producing modes submitted to /v1/videos return HTTP 400 with the canonical action endpoint in the error message.

5. Cosmos3 Parameters

Cosmos3 supports the standard SGLang video and image fields such as size, num_frames, fps, num_inference_steps, guidance_scale, negative_prompt, and seed. For distilled checkpoints, SGLang replaces num_inference_steps with the checkpoint’s fixed four-step schedule and forces guidance_scale=1.0; negative-prompt CFG and request-level flow_shift do not apply. Top-level Cosmos3 request fields:
  • max_sequence_length: maximum text token length used by the Cosmos3 tokenizer.
  • flow_shift: per-request scheduler shift for non-distilled checkpoints. If omitted, SGLang uses --flow-shift, then the mode default (3.0 for T2I, 10.0 for non-Edge video and all action modes, or 3.0 for Edge video).
  • guidance_interval: optional [start, end] noise interval for CFG. Non-distilled T2I defaults to [400, 1000]; video modes guide at every step.
Cosmos3 omnimodal fields are accepted as extra JSON fields or multipart form fields:
  • generate_sound: generate a sound track whose duration follows num_frames / fps.
  • sound_duration: explicit sound duration in seconds; takes precedence over the derived duration.
  • condition_frame_indexes: V2V latent-frame indexes to keep from the source video; defaults to [0, 1].
  • condition_video_keep: use the first or last source frames for V2V conditioning.
  • action_mode: policy, forward_dynamics, or inverse_dynamics.
  • domain_name / domain_id: select the action embodiment head.
  • raw_action_dim: number of active action dimensions; inferred for known domain names.
  • action: action array with shape [T, D], required by forward_dynamics.
  • action_fps: action-token frame rate for temporal mRoPE; defaults to the video FPS.
  • action_view_point: viewpoint used in the structured action caption.
  • action_normalization: dataset normalization mode, such as quantile, meanstd, or minmax.
Pass model-specific controls through extra_body with the OpenAI Python SDK. Raw JSON may keep them at the top level; multipart video requests should put them in the extra_params JSON object. The legacy image extra_args container remains accepted for compatibility, but new clients should use extra_body:
  • use_duration_template: whether to append SGLang’s generated duration suffix to video prompts.
  • use_resolution_template: accepted for vLLM-Omni request compatibility.
  • use_system_prompt: whether to add the Cosmos3 system prompt to the chat template.
  • guardrails or use_guardrails: per-request guardrail toggle when the server started with guardrails enabled.