Skip to main content
Prefill context parallelism (CP) distributes the tokens of a prompt across an attention CP group. Each rank computes attention and, for models with an indexer, indexer operations for its local queries, while all-gather on the KV cache makes the required keys and values available. This lets multiple GPUs share the work of processing a long prompt. Unlike tensor parallelism (TP), which partitions attention heads and weights, prefill CP partitions the token dimension. Unlike decode context parallelism (DCP), it targets prefill computation rather than dividing the decode KV-cache scan and merging partial attention results. The two features have separate configuration and support constraints.

Why prefill context parallelism?

Prefill CP offers three benefits:
  • It partitions work that would otherwise be repeated across ranks, such as indexer computation in DSA models and the DeepSeek V4 series.
  • It combines naturally with KV-cache sharding to reduce per-rank cache memory. Layer-wise sharding is available through DSA cache LayerSplit; additional sharding techniques are under development.
  • Combined with expert parallelism (EP), it can reduce communication time by using all-to-all token dispatch and combine instead of TP all-reduce. The benefit depends on the workload and interconnect.

Architecture

Strategy interface

ContextParallelStrategy separates token-layout policy from model execution and attention kernels. Server argument resolution initializes one strategy per process. Per-forward state lives in ForwardBatch.attn_cp_metadata, rather than in the model’s token-layout logic.

Zigzag strategy

For a CP group of size C, split each request’s newly extended tokens into 2C contiguous blocks. Rank r receives block r and block 2C - 1 - r. The split restarts for every request in a batch. For example, with CP=4 and 16 new tokens, each block has two tokens: Pairing an early, cheaper block with a late, more expensive block balances causal attention work while retaining contiguous query blocks. Lengths need not divide evenly: early blocks receive the remainder, and communication buffers are padded as needed. The metadata records separate query and KV lengths for the early and late blocks, including any cached prefix. Each block therefore attends to the correct causal history, not just the other tokens assigned to its rank.

Interleave strategy

For CP size C, rank r receives flattened token indices r, r+C, r+2C, .... The index runs across the batch’s newly extended tokens; it does not reset at request boundaries. Position IDs keep their original values. With CP=4 and 16 new tokens: This gives each rank queries spread across the context. For a batch containing requests of lengths 5 and 3, rank 0 owns flattened indices 0 and 4, while rank 1 owns indices 1 and 5; index 5 is the first new token of the second request. Per-request sharding metadata preserves those boundaries and each query’s causal extent.

Attention backend integration

The CP and attention-backend compatibility matrix is being expanded, with support for more combinations planned.

Compositions

Prefill CP × TP, DP, and EP

Attention TP splits heads within each CP partition; attention DP separates request batches. Dense FFNs and MoE layers can use a different layout, bridged by layer communicators. The Qwen3 example in Usage and Examples combines CP=2 with attention TP=2 and EP=4. For interleave DSA on DeepSeek V3.2 and the GLM-5 series, keep DP=1 and use the resolved attention CP=TP topology.

Prefill CP × pipeline parallelism

The eager runner preserves CP-local hidden states between pipeline stages and gathers on the final stage. PP can place stages on different nodes while keeping each DSA attention CP group within one node; it does not remove the single-machine DSA CP restriction.

Prefill CP × speculative decoding

Prefill CP can coexist with supported speculative-decoding configurations. The CP runner handles aligned speculative hidden-state inputs and gathers supported auxiliary outputs. This does not make decode or target verification context-parallel through the prefill strategy.

Prefill CP × CUDA graphs (Experimental)

The shared CP breakable-prefill-graph path currently requires zigzag, the trtllm_mha prefill backend, PP=1, and attention CP size equal to TP size. It uses CP-local static input buffers and selects capture buckets large enough for each rank’s padded input. Support for combining other attention backends with prefill CP and Breakable CUDA Graph is under development.

Prefill CP × PD disaggregation and DSA cache LayerSplit

In supported DSA deployments, --enable-dsa-cache-layer-split distributes GPU KV/indexer cache layers across CP ranks to reduce per-rank cache memory. Each rank owns a range of layers and uses scratch storage for remote-layer data. This is layer-wise cache ownership, separate from the interleave token assignment used for computation, and is not applied to draft workers. LayerSplit requires a PD prefill worker (--disaggregation-mode prefill), --enable-prefill-cp --cp-strategy interleave, and PP=1. It currently supports the mooncake and mooncake_tcp transfer backends. Do not enable LayerSplit on decode or non-PD workers; the decode side receives full cache shards through PD transfer. Use the PD disaggregation guide and model-specific launch recipes to configure both sides of the transfer; enabling prefill CP alone does not configure a PD deployment.

Usage and Examples

Use NVIDIA CUDA GPUs and select a strategy supported by your model and attention backend. The examples below target a single Linux host with Hopper GPUs and an SGLang version that includes the selected backend. CP subdivides the existing TP world; it does not add another GPU multiplier. With attention DP enabled, the attention topology is:
The server requires --tp to be divisible by --dp × --attn-cp-size. The attention layout does not by itself specify the dense FFN or MoE layout; those use their own parallelism settings.
For DSA models, such as the GLM-5 series and DeepSeek V3.2, use interleave; zigzag is currently unavailable. Do not use a smaller --attn-cp-size to request hybrid attention TP/CP for DSA models. Broader compatibility is planned for future support.

Zigzag example: GQA with attention TP and CP

This Qwen3 configuration uses four Hopper GPUs, with attention TP=2, CP=2, and EP=4 for the MoE layers. The model is a Hugging Face repository ID passed to --model-path.

Interleave example: DSA

This GLM-5.2 configuration uses eight H200 GPUs with attention CP=8 and attention TP=1. The DSA backend is selected automatically for this model; the command makes that choice explicit.
For model-specific deployment combinations, see context parallelism in the GLM-5.3 cookbook.

References