Why prefill context parallelism?
Prefill CP offers three benefits:- It partitions work that would otherwise be repeated across ranks, such as indexer computation in DSA models and the DeepSeek V4 series.
- It combines naturally with KV-cache sharding to reduce per-rank cache memory. Layer-wise sharding is available through DSA cache LayerSplit; additional sharding techniques are under development.
- Combined with expert parallelism (EP), it can reduce communication time by using all-to-all token dispatch and combine instead of TP all-reduce. The benefit depends on the workload and interconnect.
Architecture
Strategy interface
ContextParallelStrategy separates token-layout policy from model execution and attention kernels. Server argument resolution initializes one strategy per process. Per-forward state lives in ForwardBatch.attn_cp_metadata, rather than in the model’s token-layout logic.
Zigzag strategy
For a CP group of sizeC, split each request’s newly extended tokens into 2C contiguous blocks. Rank r receives block r and block 2C - 1 - r. The split restarts for every request in a batch.
For example, with CP=4 and 16 new tokens, each block has two tokens:
Pairing an early, cheaper block with a late, more expensive block balances causal attention work while retaining contiguous query blocks. Lengths need not divide evenly: early blocks receive the remainder, and communication buffers are padded as needed.
The metadata records separate query and KV lengths for the early and late blocks, including any cached prefix. Each block therefore attends to the correct causal history, not just the other tokens assigned to its rank.
Interleave strategy
For CP sizeC, rank r receives flattened token indices r, r+C, r+2C, .... The index runs across the batch’s newly extended tokens; it does not reset at request boundaries. Position IDs keep their original values.
With CP=4 and 16 new tokens:
This gives each rank queries spread across the context. For a batch containing requests of lengths 5 and 3, rank 0 owns flattened indices 0 and 4, while rank 1 owns indices 1 and 5; index 5 is the first new token of the second request. Per-request sharding metadata preserves those boundaries and each query’s causal extent.
Attention backend integration
The CP and attention-backend compatibility matrix is being expanded, with support for more combinations planned.
Compositions
Prefill CP × TP, DP, and EP
Attention TP splits heads within each CP partition; attention DP separates request batches. Dense FFNs and MoE layers can use a different layout, bridged by layer communicators. The Qwen3 example in Usage and Examples combines CP=2 with attention TP=2 and EP=4. For interleave DSA on DeepSeek V3.2 and the GLM-5 series, keep DP=1 and use the resolved attention CP=TP topology.Prefill CP × pipeline parallelism
The eager runner preserves CP-local hidden states between pipeline stages and gathers on the final stage. PP can place stages on different nodes while keeping each DSA attention CP group within one node; it does not remove the single-machine DSA CP restriction.Prefill CP × speculative decoding
Prefill CP can coexist with supported speculative-decoding configurations. The CP runner handles aligned speculative hidden-state inputs and gathers supported auxiliary outputs. This does not make decode or target verification context-parallel through the prefill strategy.Prefill CP × CUDA graphs (Experimental)
The shared CP breakable-prefill-graph path currently requireszigzag, the trtllm_mha prefill backend, PP=1, and attention CP size equal to TP size. It uses CP-local static input buffers and selects capture buckets large enough for each rank’s padded input. Support for combining other attention backends with prefill CP and Breakable CUDA Graph is under development.
Prefill CP × PD disaggregation and DSA cache LayerSplit
In supported DSA deployments,--enable-dsa-cache-layer-split distributes GPU KV/indexer cache layers across CP ranks to reduce per-rank cache memory. Each rank owns a range of layers and uses scratch storage for remote-layer data. This is layer-wise cache ownership, separate from the interleave token assignment used for computation, and is not applied to draft workers.
LayerSplit requires a PD prefill worker (--disaggregation-mode prefill), --enable-prefill-cp --cp-strategy interleave, and PP=1. It currently supports the mooncake and mooncake_tcp transfer backends. Do not enable LayerSplit on decode or non-PD workers; the decode side receives full cache shards through PD transfer. Use the PD disaggregation guide and model-specific launch recipes to configure both sides of the transfer; enabling prefill CP alone does not configure a PD deployment.
Usage and Examples
Use NVIDIA CUDA GPUs and select a strategy supported by your model and attention backend. The examples below target a single Linux host with Hopper GPUs and an SGLang version that includes the selected backend.
CP subdivides the existing TP world; it does not add another GPU multiplier. With attention DP enabled, the attention topology is:
--tp to be divisible by --dp × --attn-cp-size. The attention layout does not by itself specify the dense FFN or MoE layout; those use their own parallelism settings.
Zigzag example: GQA with attention TP and CP
This Qwen3 configuration uses four Hopper GPUs, with attention TP=2, CP=2, and EP=4 for the MoE layers. The model is a Hugging Face repository ID passed to--model-path.
Interleave example: DSA
This GLM-5.2 configuration uses eight H200 GPUs with attention CP=8 and attention TP=1. The DSA backend is selected automatically for this model; the command makes that choice explicit.References
- CP strategies and metadata: strategy interface, zigzag/interleave layouts, padding, and graph helpers.
- Prefill Context Parallelism Roadmap.
- Prefill Context Parallelism Refactor Design.
