RL Long Tail: Mid-Batch Parallel Shape Switching

On four A100s with no NVLink, switching an RL rollout from DP4 to TP4 once the long tail remains is 1.31 times slower than staying on DP2×TP2. The two-GPU custom all-reduce keeps the mixed shape ahead in both the crowded phase and the sparse phase, and one switch costs at least 19 seconds, more than it can save.

Training a reasoning model with reinforcement learning starts each step by letting the current policy write an answer to every problem in a batch. That step is the rollout. Answer lengths spread widely. Most finish in a few thousand tokens. A few write until the length cap. The first half of the batch keeps every GPU full. The second half is a handful of long requests, and the batch cannot finish until they do. A natural guess is to serve the crowded half in a high-throughput shape, then switch to a low-latency shape for the tail. This article tests that guess on an ordinary four-GPU PCIe machine. It does not hold. The reason is more useful than the guess.

Three ways to place the same four GPUs

An 8B dense model in bf16 is about 16 GB, and one 40 GB A100 holds it. Data parallelism, DP, puts a full copy on every GPU and serves a separate slice of the requests. Four GPUs are DP4, four queues and no communication between them. Tensor parallelism, TP, splits each layer’s weight matrices across GPUs, the Megatron-LM cut (Shoeybi et al. 2019). After the layer, an all-reduce adds the partial results. Four GPUs as one engine are TP4. Between them is a mixed shape, DP2×TP2, two TP2 engines of two GPUs each.

The tradeoff sits in decode. Each step emits one token per request in the batch, and it has to read the whole weight tensor plus the KV cache of the context so far. A100 memory bandwidth is about 1.555 TB/s, so reading 16 GB of weights is about 10 ms. That is the order of a single-GPU, single-request token time. TP cuts the weights into N pieces, so each GPU reads 1/N, and the token time falls. The cost is two all-reduces per layer. This Llama-shaped 8B model has 32 layers, so a step does 64 all-reduces. The message is the batch size times hidden size 4096 times 2 bytes. At 16 requests that is 128 KiB. DP has no communication, and each GPU reads every weight, so a single request is slower. Its advantage is throughput, because four engines can each grow their own batch. The usual rule of thumb is that DP wins when many requests are in flight, and TP wins when few remain.

In a batch of 16 requests at about ten thousand tokens of context, each request’s KV cache is about 128 KiB per token: 32 layers, 8 KV heads, head dimension 128, key and value, bf16. The batch is about 20 GB, larger than the weights. KV bytes per GPU depend on how many requests that GPU serves. At the same total number of requests in flight, TP4 and two TP2 engines read the same KV bytes per GPU. Only the weight read differs.

Which path the all-reduce takes

Whether the all-reduce is cheap depends on the path. NCCL’s ring cuts the data into chunks and passes them 2(N−1) times around N GPUs (NVIDIA, NCCL). It is bandwidth-optimal for large messages, and a small message pays tens of microseconds of protocol. Decode all-reduces are many small messages, so vLLM (Kwon et al. 2023) ships its own custom all-reduce. GPUs map each other’s buffers with CUDA IPC and reduce in one kernel over peer-to-peer reads, either one-shot or two-shot, and skip the NCCL protocol (vLLM custom all-reduce).

The implementation turns on only for a full NVLink mesh, or for exactly two PCIe-only GPUs. With more than two PCIe-only GPUs, vLLM logs that custom all-reduce is disabled and falls back to NCCL. On this machine two A100s share a PCIe switch, and the two switches meet at the CPU host bridge. A measured 128 KiB all-reduce is about 20.6 µs on the two-GPU custom path and about 93.4 µs on the four-GPU NCCL path, more than four times slower. Times 64 per step, that is about 1.3 ms against about 6 ms, a first-order term inside a decode step of ten-plus milliseconds.

Three placements of four A100s. DP4 has no all-reduce. DP2×TP2 keeps each all-reduce under one PCIe switch on the custom path. TP4 crosses the host bridge on an NCCL ring, and custom all-reduce is off.
Figure 1. The same four A100s. (a) DP4 is four independent queues and no all-reduce, and every GPU reads the full weights. (b) DP2×TP2 keeps each TP pair under one PCIe switch, on vLLM's two-GPU custom all-reduce. (c) The TP4 ring crosses the host bridge, and custom all-reduce is off.

Going from TP2 to TP4 is not one more GPU on the same path. It is a different communication implementation and a longer physical route.

The long tail, and systems that already reconfigure

RL stacks usually run the rollout on an inference engine. HybridFlow, verl, alternates generation and training on the same GPUs (Sheng et al. 2024). Reasoning training makes rollouts long. DeepSeek-R1 reports that answer length keeps growing during training (DeepSeek-AI 2025). Kimi k1.5 cuts over-long trajectories and resumes them across iterations (Kimi Team 2025). The batch here has the same shape. Problems come from MATH-500 (Lightman et al. 2023), itself drawn from MATH (Hendrycks et al. 2021). The median output is about 2400 tokens, the 90th percentile about 9800, and about 3% to 4% of requests write to the 16K cap.

Systems already change deployment shape at runtime. Llumnix migrates live requests across instances (Sun et al. 2024). DistServe (Zhong et al. 2024) and Splitwise (Patel et al. 2024) put prefill and decode on different GPUs. NVIDIA Dynamo decides per request whether to split them, from the effective input length and the prefill load (NVIDIA Dynamo). That choice already switches at zero migration cost. TP degree has no in-process switch. The available pattern is one resident engine per shape. An idle engine enters vLLM sleep mode, swaps weights to host memory, frees the KV pool, and wakes in a few seconds (vLLM sleep mode). A request that must change shape is cancelled and resubmitted on the target engine, with the problem plus the tokens already written as the new prompt. The cost is a prefill of that context.

TL;DR

The question is whether a long-tailed RL rollout finishes sooner if the parallel shape changes mid-batch. The guess is that DP4 is the high-throughput first half and TP4 is the fast-per-token second half, so one switch takes both.

On four A100s in two PCIe-switch pairs, the switch point chosen from an offline profile is 1.31 times slower than static DP2×TP2. DP2×TP2 loses neither phase. TP4 is only 0% to 11% faster per token than TP2 in the tail. One switch pauses for at least 19 seconds.

The makespan of a long-tail batch is about the length cap times the lifetime average token time of the longest request. On a PCIe-only machine, that time is set by whether the all-reduce can stay under one switch.

The hypothesis

There is a threshold n*, the number of requests still in flight. Run DP4 until only n* requests remain, move them to TP4, and finish. The makespan, the time of the last completion, should beat the best of DP4, TP4, DP2×TP2, and a length-based spatial split by at least 15%. The paired 95% confidence interval should exclude zero. Every switch and recompute cost counts. The 15% bar is there so a tuning delta does not pass.

The guess looked plausible from a single-request microbenchmark and an order-of-magnitude estimate. At one request and 1K of context, four-GPU TP4 is about 2.4 times faster per token than one GPU, and at a loose latency target DP4’s capacity is about 2.5 times TP4. A two-phase closed form says why both sides have to be large. When the two phases each take half the time, picking the best shape per phase beats the best static choice by 15% only if each side leads by about 36% in the symmetric case. A large lead on one side and a small lead on the other gives almost nothing. Putting the curves into a fluid simulation of 128 requests with log-normal lengths, the most favorable cell is four GPUs, a wide length spread, and a 16K cap. The median gain there is 5.9%, and the maximum is 27%. The estimate’s weakness was written down at the time. The curves are from 1K of context, a long context reads more KV per step, and the comparison is only TP4 against DP4. The mixed shape is not in it.

How it was measured

The machine is one node, four A100-PCIe 40 GB GPUs, two per PCIe switch, the switches joined by the host bridge, and no NVLink. The software is vLLM 0.25.1, PyTorch 2.11, CUDA 13.0, bf16, with prefix caching, chunked prefill, and CUDA graphs at their defaults. The model is DeepSeek-R1-Distill-Llama-8B. The workload is the first 256 problems of MATH-500 after a fixed shuffle, with the R1 math prompt, temperature 0.6, top_p 0.95, and a cap of 16384 tokens, ending on EOS. All 256 requests arrive at time zero. The other 244 problems are only for the offline profile. Three seeds. Within a seed every configuration sees the same problems and the same per-request sampling seed. Statistics are paired by seed. The 95% interval uses a t distribution with 2 degrees of freedom.

Five families. DP4, TP4, and DP2×TP2 are static. SPLIT@L is a spatial split. New requests enter two TP1 engines on two GPUs, and a request whose output passes L tokens moves to a TP2 engine on the other two GPUs. L is 2048, 4096, or 8192. DYN@n is the dynamic policy. Start on DP4. When unfinished requests fall to n, cancel them, wait for the four TP1 engines to drain, put those engines to sleep, wake a prewarmed TP4 engine, and resubmit the remainder with the tokens already generated. n is 8, 16, 32, 64, or 128. n=32 is the point an offline predictor, looking only at the profile batch, picked before the measurement batch ran. That point was registered in advance. Every configuration runs on the same resident engines, and unused engines are in L1 sleep. Standalone DP4 and TP4, not coresident and not using sleep, differ from the coresident versions by 1.0% to 1.2%. Output token count and answer accuracy are controls, so an early stop cannot fake a speedup.

One measurement trap is worth stating. transformers 5.x loads this model’s tokenizer as a Llama-2-style class and eats spaces on encode. The server tokenizer in the same vLLM image does the same. The model still answers, and nothing errors. The first smoke run was discarded. The reported runs encode with tokenizer.json directly and check that decoding returns the original text.

The mixed shape wins both phases

Batch makespan for static, dynamic, and spatial-split shapes, and per-request token time against the number of requests in flight. DP2×TP2 is the shortest. The predicted dynamic switch is 1.31 times longer.
Figure 2. (a) Makespan, mean and 95% interval, as a multiple of DP2×TP2. Hatched bars are the dynamic switches. Cross-hatched bars are the spatial splits. DP4-sa is the standalone DP4. (b) Per-request token time against requests in flight on the four GPUs.

DP2×TP2 finishes in 304 seconds. The preregistered DYN@32 takes 399 seconds, 1.31 times DP2×TP2. The paired difference is +94.9 seconds, 95% interval [+8.9, +181.0], and it excludes zero. TP4 is 411 seconds. DP4 is 452 seconds. The three spatial splits are 493 to 545 seconds. A one-seed standalone TP4 at 415 seconds is not on the figure. The dynamic policy is 11.6% faster than DP4, interval [0.9%, 22.2%], and its gap to TP4 is inside the noise. The reversal between DP4 and TP4 is real, and a switch collects about 10% of it. A mixed shape that wins both phases takes that gain away. The paired difference in output tokens is +25K, inside the seed noise (standard deviation 23K). The accuracy difference is −0.9 percentage points, under twice that noise.

The ordering is the request that finishes last. In every configuration and every seed, that request hits the 16384 cap. Makespan is almost 16384 times its lifetime average token time. On DP2×TP2 that average is 18.5 ms, which is 303 seconds. On TP4 it is 24.9 ms, which is 408 seconds. Both match the measurement. A long request lives through a crowded phase, hundreds of requests together, and then a sparse phase of a few requests. Makespan is the weighted average of those two phases, not the tail speed alone.

Figure 2(b) splits the weight. The horizontal axis is requests in flight on the four GPUs. The vertical axis is the token time a request actually sees, engine concurrency times the sample interval divided by tokens that engine emitted, median across the three seeds. In the crowded phase TP4 is worst. At 256 in flight it is about 49 ms per token, twice DP2×TP2 at about 25 ms, because one TP4 engine pushes 256 requests’ all-reduce messages, about 2 MiB, across the host bridge, and it has one scheduler. DP4 is close to DP2×TP2 while crowded, and it is the slowest in the sparse phase, about 23 ms, because every GPU reads the full weights. At 8 to 16 in flight, TP4 and DP2×TP2 nearly meet. At 16 they are 15.8 ms against 15.5 ms. Read per engine instead, TP4 at 16 in flight is 15.1 ms and a TP2 engine at 8 in flight is 16.9 ms, so TP4 is 11% faster. That is the reading most favorable to TP4. Neither reading has the 2.4 times the desk estimate used.

The same order-of-magnitude split explains it. At the same total number of requests in flight, TP4 and two TP2 engines read the same KV bytes per GPU. The weights differ. TP2 reads 8 GB per GPU per step, about 5.1 ms. TP4 reads 4 GB, about 2.6 ms, and saves about 2.6 ms. On the other side, 16 requests mean 64 all-reduces of 128 KiB. The two-GPU custom path is about 1 ms. The four-GPU NCCL path is about 6 ms. The two terms are the same size and opposite in sign, so the tail advantage lands between zero and about ten percent. Every shape’s step time is 2 to 5 ms above a pure bandwidth model, and TP4 is the furthest above it. This article does not break that remainder into sampling, scheduling, and the all-reduce. The conclusion depends on custom all-reduce being off for four GPUs, and on the two groups crossing a host bridge. On a full NVLink mesh, TP4 would still use the custom path or a faster one, and this arithmetic would change.

What one switch costs

Switch pause against the number of requests left, stacked as cancel, sleep, wake, and re-prefill, next to the maximum saving versus DP2×TP2. Measured makespan of a DP4-to-TP4 switch stays above static DP2×TP2 at every switch point.
Figure 3. (a) One DP4-to-TP4 switch. The stacked pause is cancel and drain, sleep, wake of TP4, and re-prefill. The green line is the most that switch can save against DP2×TP2. (b) Measured makespan against the switch point, the offline prediction, and the two static baselines.

Cancel and drain are about 0.2 seconds. Sleeping the four TP1 engines is about 2.7 seconds, a copy of four 16 GB weight sets back to pinned host memory. Waking TP4 is about 2.2 seconds. Then comes the re-prefill, from resubmit until every migrated request has its first new token. The fixed part stays at 5.0 to 5.4 seconds. Re-prefill grows from 14 seconds at n*=8 to 43 seconds at n*=128. The whole pause is 19 to 48 seconds. Re-prefill dominates for a one-line reason. At n*=32, 32 migrated requests carry about 8K of context each, about 257,000 tokens. TP4 prefill is only about 7.8K tokens per second. 32 times 8K divided by 7.8K is about 33 seconds, and the measurement is 32.8 seconds. The prefill rate itself is the less intuitive part. Four-GPU TP4 is slower than one GPU running TP1, about 11.2K tokens per second, because prefill all-reduces are large and the host-bridge bandwidth is the limit. The same topology fact both shrinks TP4’s tail gain and raises the cost of switching to it.

The green line is an upper bound on the saving against DP2×TP2. It is analysis, not a measurement. Take DP2×TP2’s own tail duration and multiply by the 11% per-step advantage that is the most favorable reading for TP4. The tail while at most 8 requests remain averages 19.9 seconds, so the bound is about 2.2 seconds. At most 16 requests, the tail is 93.5 seconds and the bound is about 10.3 seconds. Above 16 in flight, TP4 is slower than DP2×TP2, so switching earlier only loses, and the bound caps near 10 seconds. The measured pause starts from DP4. Starting from DP2×TP2 would sleep two TP2 engines. The fixed part is the same order, and the re-prefill follows the migrated context, the same order. The smallest pause, 19 seconds, is about twice the most the switch can save. No switch point makes a one-time move off the best static shape pay for itself.

Figure 3(b) is the predictor. The solid line is measured DYN@n makespan. The dashed line is the offline predictor, using only the profile batch’s DP4 and TP4 curves, a switch probe, and the prefill rate. It picks n*=32, which is also the measured best point. That agreement carries little information. The points at 32, 64, and 128 differ by 2.3%, less than the seed standard deviation of one configuration, 11 to 18 seconds. The absolute prediction is systematically 6% to 16% optimistic. The bias sits in the DP phase. Static DP4 is underestimated by 15.7%, TP4 by 1.4%. Feeding the real output lengths does not fix it, at −7.2%, because the model looks up throughput by per-engine concurrency and does not represent the imbalance after round-robin assignment. The four DP engines’ last completions differ by 90 to 127 seconds. The switch-cost model, a fixed 5.5 seconds plus context divided by the prefill rate, errs by −2% to −6%. An offline profile can rank the switch point and can price the switch. What it misses is the tail of the DP shape.

Why the conjecture does not hold

The hypothesis fails in its own experiment, and what fails is the premise. A 2.4 times faster TP4 token in the tail does not hold at long context and a dozen requests in flight. The measurement is zero to about ten percent. Without that premise, the two-phase relation never gets the roughly 36% lead it needs on both sides. The crowded-phase winner is not DP4. It is DP2×TP2. A static mixed shape wins both phases, and any single switch that pays a recompute can only be slower. The first-round estimate that picked the scenario compared only TP-N with DP-N. It left the mixed shape out, and the mixed shape is what won.

A few gaps remain. Three seeds leave a wide interval. DYN@32’s makespan interval is [356, 442] seconds. The paired difference against DP2×TP2 still excludes zero, so the direction is stable. A network job that uses no GPU was running on the same node, and its NIC shares a PCIe switch with two of the GPUs. Host CPU was 5.4% busy over the run. That interference was not removed. If it matters, it is more likely to hurt the cross-switch TP4 path, which is the direction against the dynamic policy, and it is not large enough to explain a 95 second gap. One model, one length cap, and one topology were measured. A shorter cap makes the tail shorter and gives TP4 less time to help, which is worse for switching, and it was not measured. Splitting or fusing prefill and decode was not measured. Per-request choices of that kind already exist in shipping systems at zero switch cost, and this article does not judge their gain.

What sets the makespan

The makespan of a long-tail batch is about the length cap times the lifetime average token time of the longest request. In every configuration and every seed the last request hits the 16K cap. 18.5 ms times 16384 is 303 seconds, and 24.9 ms times 16384 is 408 seconds, within 1% of the measured makespan. What to optimize is a shape’s weighted behavior in the crowded phase and the sparse phase, not its peak tail speed. This holds when output length has a hard cap and some requests hit it. About 3% to 4% of this batch do.

On a PCIe-only multi-GPU node, the best parallel shape is set by the board topology. A concurrency-1 microbenchmark says the opposite. vLLM 0.25.1 enables custom all-reduce only for two PCIe-only GPUs. With four A100s on two switches, TP4 falls back to NCCL across the host bridge. In the tail it is only 0% to 11% faster per token than TP2. In the crowded phase its aggregate throughput is about half of DP2×TP2, and its prefill is slower than one GPU. DP2×TP2, with each TP pair under one switch, wins both phases. Its makespan is 26% shorter than TP4 and 33% shorter than DP4. This is for no NVLink, two GPUs per switch, an 8B dense model, and contexts of thousands to tens of thousands of tokens. A full NVLink mesh is a different machine.

A dynamic choice between two shapes pays only when the winner of each phase leads by about a third. When the two phases take similar time and the switch is free, each side needs about a 36% lead before phase-by-phase selection beats the best static choice by 15%. A large lead on one side and a small lead on the other goes to about zero. Charging the switch only raises the bar.

The pause of moving a live request across parallel shapes is a fixed few seconds plus the migrated context divided by the target shape’s prefill rate. The saving cannot exceed the target’s per-step advantage times the remaining tail. Both can be computed offline before the switch. Here the fixed part is 5.0 to 5.4 seconds, sleep 2.7 and wake 2.2. Re-prefill is 70% to 90% of the pause. The cost model errs by −2% to −6%. The smallest pause is 19 seconds against a maximum saving of about 10 seconds. This assumes a prewarmed engine per shape, continuation by recomputing the tokens already written, and one PCIe node. A primitive that moves KV across layouts would shrink the recompute term. None was available to measure.

A comparison of parallel shapes has to include the mixed shape. TP-N against DP-N shows a real reversal. The switch is 11.6% faster than DP4, and the interval excludes zero. That reversal does not change the deployment, because DP2×TP2 is faster than both ends. This applies when the GPU count allows a mix, four or more, and the topology groups the GPUs into better-connected pairs.

These numbers cover one machine of four A100-PCIe 40 GB GPUs, two PCIe-switch pairs, no NVLink, one 8B dense reasoning model, one math-rollout batch, a 16K cap, and three seeds. They do not say that a structural dynamic switch is never worth it. They do not say DP2×TP2 dominates on NVLink, on another PCIe topology, or on a larger model. They say nothing about a prefill-decode phase reversal. A −0.9 percentage point accuracy gap is not evidence that the dynamic switch hurts or helps answer quality.

The mixed shape already covers both phases

DP2×TP2 is ahead while the batch is crowded and it is not behind in the tail. Moving to TP4 for the tail pays at least 19 seconds to save at most about 10. On this board, the shape that keeps each all-reduce under one switch is the one to run for the whole batch.

Where Communicator Rebuild Actually Costs measures the NCCL cost of rebuilding a communicator when the parallel shape changes. When the shapes can be listed in advance, a prewarmed pool already takes that cost off the critical path.

Hybrid PD Handoff: Does Waiting by Component Beat Waiting by Layer? asks, once prefill and decode are already on different GPUs, whether the handoff should wait per component or per layer. On that model, the extra gain of waiting per component stays near zero.