Table of Contents A single vllm serve process does three jobs that get in each other's way: processing prompts (prefill), generating tokens (decode) and a pile of CPU work around them. Disaggregated serving in vLLM separates the different stages of LLM inference. Splitting prefill from decode stops long prompts stalling everyone else's output, as long as the KV cache moves between them fast. Moving tokenization and parsing to a CPU-only frontend ( /render , /derender ) takes that work off your GPU nodes and leaves the engine working purely in token IDs. This post covers when each split or separation is worth using, how to run it with vLLM v0.30.0 or later, and the key gaps and improvements still ahead. One Server Doing Three Unrelated Jobs Start a plain vllm serve and you get one process handling three workloads that have nothing in common. One long prompt arrives while dozens of requests are mid-response and every one of those streams stutters until it's processed. It's the spike you see the moment concurrency goes up. The cause is two phases with little in common sharing one GPU. Prefill reads the whole prompt in one pass, is limited by compute and sets your time to first token (TTFT). Decode is the opposite. It produces one token at a time and what holds it back is how fast the GPU can read the model weights from memory which sets your inter-token latency (ITL). Put them on the same GPU and every decode stream waits while a long prefill runs. The third job however is quieter and easier to miss. Chat templating, tokenization, detokenization, reasoning parsing, tool call parsing. All pure CPU work running on a box you're renting for its accelerators. Disaggregation is the obvious response: stop making one process do all of it. If you want the vLLM engine level picture before you go further, Inside vLLM covers how the scheduler mixes prefill and decode today and has a short section on P/D. The Dimensions vLLM offers several points where work can be split and they can be combined. Prefill and decode run as two instances. What passes between them is the KV cache: the attention keys and values for every prompt token which decode needs before it can emit anything. It's big. Llama-3.1-70B in BF16 stores 320 KiB per token, so a 10k-token prompt hands decode about 3 GB. That's roughly 65 ms at a 400 Gb/s line rate before any overhead and all of it lands on TTFT. A KV connector moves it, usually over RDMA. There are more than a dozen connectors upstream now, including NIXL, LMCache, Mooncake, FlexKV and AMD's MoRI-IO, plus a MultiConnector that chains them. The frontend comes off the GPU box entirely. /render turns an OpenAI request into token IDs, the engine runs token-in / token-out and /derender turns the output token IDs back into a proper OpenAI response with content , reasoning and tool_calls split out. That last leg only landed recently and it completes the round trip. Beyond P/D , there's encoder disaggregation for multimodal and the AFD plugin for splitting attention from FFN in MoE models. Both use the same idea at a different point in the model. P/D itself now also covers hybrid SSM models with Mamba state transfer included. Collocated vLLM serving versus a four-tier disaggregated pipeline Figure 1. One process doing everything, versus the same pipeline cut into four tiers. What It Buys You and What It Costs The pitch isn't peak throughput. With no latency target, splitting the same GPUs into prefill and decode pools won't necessarily move more tokens per second. What it buys is goodput : the request rate you can sustain while requests still meet both their TTFT and ITL targets. You can tune TTFT and ITL independently. Different parallelism on each tier, sized for the phase it's running. Prefill can be TP-heavy, decode can be sized for batch. Neither change drags the other one with it. Tail latency stays low as load climbs. A decode instance that only runs decode batches never has a long prefill stall its streams. Chunked prefill gets you partway there, but the right chunk size depends on the traffic, so you end up retuning it. This was measured on one box with two NVIDIA L40S GPUs (48 GB, PCIe, no NVLink): Qwen2.5-7B-Instruct, ~8k-token prompts, 256 output tokens, 100 Poisson-arrival requests per rate, prefix caching off. Collocated is one vllm serve --data-parallel-size 2 , so both setups get the same two GPUs. That isn't the strongest collocated baseline because DP ranks can hold each other up and two independent replicas behind a load balancer might do a little better. P/D is one prefiller and one decoder over NIXL behind the example proxy. p99 and median inter-token latency against offered load from 0.2 to 2 req/s. Collocated p99 jumps from 23 ms to 169 ms at 0.4 req/s and reaches 263 ms at 2 req/s. P/D p99 stays between 25 and 52 ms. Figure 2. Median ITL is 21–24 ms in both setups up to 1 req/s. At 0.4 req/s, collocated p99 jumps to 169 ms while P/D holds at 29 ms and never exceeds 52 ms. At 0.4–0.6 req/s the medians match and the collocated p99 is about six times higher. That's 8k-token prefills landing on a GPU that's also decoding and stalling every stream on it until they finish. P/D keeps them off the decode GPU entirely. Goodput goes up when the transfer is fast. AMD's single-node MoRI-IO benchmark ran Qwen3-235B-A22B-FP8 at 8 req/s on one 8-GPU MI300X node. 73 of 100 requests met both a 1 s TTFT and a 50 ms ITL target, against 30 of 100 for collocated serving. That's about 2.4× the goodput on the same hardware. At cluster scale, llm-d's P/D guide reports about 59% lower mean end-to-end latency and 67% lower P95 for gpt-oss-120b on 16 H200s, compared with the same GPUs run as aggregated replicas. Every first token pays for the transfer. Those results assume the KV cache moves fast: RDMA over InfiniBand or RoCE between nodes, NVLink or GPU peer-to-peer within one. Our test box had neither. The L40S pair can't do peer-to-peer copies ( nvidia-smi topo -p2p r reports NS ) and each 8k prompt's ~470 MB of KV cache (56 KiB per token for Qwen2.5-7B) took about 1.3 s to reach decode. At 0.2 req/s, P/D median TTFT was 2.2 s against 0.7 s collocated, so it missed a 2 s TTFT target at every rate even though its ITL tail stayed flat. Bandwidth alone doesn't explain 1.3 s. Even staged through host memory, PCIe 4.0 should move that in tens of milliseconds. Most of it is overhead around the copy, including decode only noticing a finished transfer when it polls between its own forward steps. So check the transfer before you benchmark anything else. On one box, nvidia-smi topo -p2p r should say OK between your prefill and decode GPUs. Then send a few long prompts one at a time and read decode's KV Transfer metrics line. If Avg xfer time is in the hundreds of milliseconds, fix that first. The CPU tier is cheap. Once tokenization and parsing move off the GPU box, you scale them against CPU load instead of buying accelerator time to run a tokenizer. Long prompts, multimodal preprocessing and reasoning or tool parsing are where the render tier does real work and none of it needs a GPU. On the same box, templating and tokenizing a 9k-token chat prompt for Qwen2.5-7B cost about 15 ms of CPU. One render server with default settings topped out at 73 req/s using just over one core. At the 0.4 req/s where collocated's tail fell apart, rendering is under 1% of one core. The catch: you now operate three or four services instead of one and the KV transfer is a new failure mode. Collocated is still the right answer for plenty of deployments. Your situation Recommendation ITL p99 misses your SLO under production load Disaggregate. This is the main use case. Long prompts at high concurrency Disaggregate, if your KV transfer is fast. Prefill interference is worst here. Chat or agent loops over a growing context Disaggregate, with bidirectional transfer (below). Pair it with KV offloading or a shared KV cache such as LMCache or Mooncake. Templating, tokenizing or parsing shows up in your GPU nodes' CPU profile Split off the render tier. TTFT is the binding constraint Stay collocated or measure first. The transfer lands on every first token. Your KV transfer is slow (check decode's KV Transfer metrics ) Fix it or stay collocated. It lands on TTFT and caps throughput. Check the fabric too, since a misconfigured network usually slows collectives as well. Low, bursty or latency-insensitive traffic Stay collocated. Running Prefill/Decode Everything from here on assumes vLLM v0.30.0 or later. The examples use Qwen3-0.6B because it loads fast. That's fine for checking the wiring but it's too small to show a P/D benefit, so benchmark with a larger model (see Where to start ). Three processes: prefiller, decoder, proxy. # Prefiller on GPU 0 CUDA_VISIBLE_DEVICES = 0 UCX_NET_DEVICES = all VLLM_NIXL_SIDE_CHANNEL_PORT = 5600 \ vl
Table of Contents A single vllm serve process does three jobs that get
Table of Contents A single vllm serve process does three jobs that get in each other's way: processing prompts (prefill), generating tokens (decode) and a pile of CPU work around them. Disaggregated serving in vLLM separates the different stages of LLM in
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。