1. Model Introduction
FLUX 3 Action is a robot policy from Black Forest Labs built on a FLUX 3 video diffusion transformer. From the current camera frames, the robot state and a language instruction, it denoises a chunk of future video latents jointly with a chunk of continuous actions; only the actions are returned. The model has three parts:- DiT (
JointSingleSeq, about 6.6B parameters). Each stream (text,video,video_cond, the action and state streams) runs through its own 5 mode blocks. The text and all streams then share 28 joint single-stream blocks with per-stream modulation and a 4-axis(t, h, w, l)RoPE. - Text encoder: Qwen3-VL-4B. Eight hidden layers are stacked into a 20480-dim context.
- Video VAE: a Swin3D neighborhood-attention VAE (NATTEN), 96 latent channels, 32x spatial compression.
4.0 on video, 1.0 on actions).
SGLang runs it in the native multimodal_gen runtime. Work that does not depend on the noised streams is computed once:
- The text context, once per caption, across requests.
- The observation and state streams, once per request.
- The mode blocks of the noised streams, once per step, shared by the conditional and unconditional passes.
black-forest-labs/flux-3-action-droid; three 360x640 cameras wrist, left, right; state and action dim 8; 32-action chunks):
Latency is the median server-side time per request (
server_timing.infer_ms) over 50 sequential requests after 10 warmup requests, on one GPU with the default serve command, eager mode (no torch.compile). Requests go through the OpenPI WebSocket with msgpack numpy images, and the caption is cached. Measure it with python -m sglang.multimodal_gen.benchmarks.bench_flux3_action --url ws://127.0.0.1:30000 against a running server. On H200 the FP8r packages are slower than BF16; they save memory. LayerNorm + modulation, SwiGLU and the gated residual run on bit-exact fused kernels; QK RMSNorm + RoPE runs on a fused kernel at bf16 rounding level (set SGLANG_ENABLE_FUSED_QKNORM_ROPE=0 to use the eager path). FP8r packages load their native E4M3 weights (one scale per output row) and run torch._scaled_mm with per-token activation scales. This needs an SM89+ GPU.
The frozen text encoder and video VAE are downloaded from the pinned revision of black-forest-labs/flux-3-action-base that the policy config references.
References:
2. Installation
Command
sglang[diffusion]; without it the VAE uses a compiled FlexAttention fallback with the same windows (actions stay at the bf16 rounding level). The fallback compiles on the first request (about 3 s on H200) and is then no slower than NATTEN for this single-frame encode (24 ms vs. 36 ms per request on H200). To use NATTEN, pick the wheel that matches your torch and CUDA versions from whl.natten.org.
3. Model Deployment
Serve the DROID policy:Command
Command
manifest.json, config.native.json and model.safetensors) is detected automatically when passed as the model path. Pass --revision to pin a Hub revision. At startup the policy config and weights are checked against the SHA-256 hashes in manifest.json.
Peak GPU memory and per-request latency on one RTX 5090. The resident rows use the same median protocol as the table above. The layerwise-offload rows are from the earlier single-run measurement.
Layerwise offload streams the DiT blocks (and, with
all, the text encoder and VAE blocks) from host memory, so use it only when the resident configuration does not fit.
3.1 Multi-GPU
Three layouts split the DiT across GPUs. They compose as--num-gpus = TP size x SP degree x (2 with CFG parallel):
CFG parallel only helps recipes that use guidance (
base, fp8r). The distilled gd and sd variants run one pass per step, so both GPUs compute the same pass. TP and SP accept the FP8r packages too. Ring attention is not supported.
Command
3.2 Action Request Schema
The response returns absolute commands of shape
[32, 8] in the dataset’s conventions.
4. API Usage
4.1 Generic Action HTTP API
Example
GET /v1/actions/metadata reports the camera keys, action shape and sampler defaults of the served policy.
4.2 OpenPI-Compatible WebSocket
/openpi/policy takes one observation per message, with camera images as
observation.images.<camera> (wrist, left, right, or the LeRobot names
wrist_image_left, exterior_image_1_left, exterior_image_2_left), the state
as observation.state and the instruction as task or prompt. The response
carries the [32, 8] chunk as actions. Use the msgpack helpers from the
Pi0.5 page to pack numpy arrays.
5. Accuracy
With the same observation and seed, SGLang matches the FLUX Action reference implementation to the bf16 rounding level:- The DiT agrees to 1e-6 (relative) in fp32.
- The VAE latents and text contexts are bit-identical.
- Across the full 4-step, CFG 4.0 sampling loop, actions differ by at most 0.015 rad (mean 0.003 rad). The reference’s own eager and prepared paths differ from each other by 0.015 rad.
- FP8r packages differ from the reference FP8r path by at most 0.025 rad. The reference FP8r path itself differs from BF16 by 0.055 rad.
