Where Communicator Rebuild Actually Costs

Rebuilding an NCCL communicator on demand takes 50 to 180 ms, and almost all of that is the lazy connect inside the first collective. Warming a pool that uses about 1% of GPU memory moves that cost off the critical path.

Parallel shapes are starting to change at runtime

Large-model inference and training used to assume one thing. From the start of a distributed run to the end, the set of GPUs does not change. The degree of tensor parallelism, which splits one layer’s matrix by columns or by rows and merges with an allreduce at the end of the layer, is fixed at launch and then left alone. That assumption is loosening. Seesaw switches the parallel strategy between prefill and decode, because the two phases have very different compute-to-memory ratios and one degree cannot be best for both (Su et al., 2025). LoongServe lets the degree of sequence parallelism grow and shrink with request length (Wu et al., 2024). Llumnix migrates requests between instances at runtime and reports a pause of only tens of milliseconds (Sun et al., 2024). On the training side, torchft rebuilds the communicator from each step’s quorum when membership changes (PyTorch, torchft).

Once the set of participants can change, the communication layer becomes the problem. A collective library organizes communication around a communicator. A group of processes agrees on a unique id, each joins under a rank, and the library builds ring and tree topologies, allocates channel buffers, and opens intra-node P2P and inter-node InfiniBand connections. Rank is both the communication identity and the physical place. When membership changes, the old communicator no longer describes reality. The standard response today is to drain in-flight work, destroy the old group, and init a new one.

A few scales are worth holding onto. A 1 MB allreduce on four A40s is under a millisecond. 64 MB is about 3 to 7 ms. A decode step is usually tens of milliseconds. In an earlier experiment, a cross-GPU request migration paused for a median of 69 ms and happened about once every 2.6 s. If rebuilding a communicator takes hundreds of milliseconds, it is a first-order term in a single event.

What NCCL already offers for a changing group

Over the last few releases NCCL has been turning the communicator from a one-shot object into a mutable one. ncclCommSplit cuts a subgroup out of an existing communicator by color. With splitShare, the child can share the parent’s transport resources, which is meant to avoid reconnecting (NVIDIA, NCCL API). ncclCommShrink drops ranks from an existing communicator. With NCCL_SHRINK_ABORT it first stops in-flight work and then shrinks. ncclCommRevoke stops every in-flight operation and waits until the communicator is quiet, after which it can be destroyed, split, or shrunk. Since 2.29, ncclCommGrow adds ranks to an existing communicator, and together with shrink it follows nodes as they fail and return (NVIDIA, NCCL 2.29.2 release notes). ncclCommWindowRegister registers a buffer on a communicator for symmetric-memory kernels and the device API. The window is bound to that communicator, so a new communicator means a new registration.

These primitives avoid a full teardown when membership changes. They share a boundary. None of them keeps ordered in-flight messages. Revoke and SHRINK_ABORT drop work in progress. After a shrink, ranks are renumbered. The new communicator is a new object, and operations already queued on the old one do not move. PyTorch c10d wraps this with split_group, shrink_group, and new_group, and by default it creates the NCCL communicator lazily, on the first collective (PyTorch, torch.distributed).

There is also an old approach that is not an API. If the target shapes can be enumerated in advance, build a communicator for each of them at startup and switch handles. On a single node of 4 to 8 GPUs there are only a few reachable TP shapes. That path has always existed on paper. Nobody had priced it.

TL;DR

Inference and training are starting to change the parallel shape, and the set of GPUs, while a job is running. NCCL binds a communicator to a rank, so a new shape means tearing the old group down and building a new one. A natural guess is to detach identity from rank, and then a migration or a TP reshape would not need that teardown.

On A40 and V100, on one node and across InfiniBand, NCCL 2.20 through 2.32, an on-demand rebuild takes 50 to 180 ms. A request migration pauses for 69 ms, so the two are the same order. Almost all of the rebuild is the lazy connect inside the first collective. Split, shrink, and grow are not much cheaper than a fresh init. That connect can be paid at startup.

Prebuilding TP4, TP2, and TP1 on 4 GPUs costs an extra 464 MB per GPU, about 1% of memory. Collective bandwidth does not drop, and the switch is sub-millisecond. Abort and revoke sit on a floor of about 0.5 s, and that can run in the background. When the shapes can be listed in advance, the primitives that already exist have already removed this cost. A rank-independent identity does not buy anything further.

The conjecture we can test

The gap we thought we saw was the lack of a layer that separates a logical communication identity from a physical rank. If a logical flow has a stable endpoint, in-flight work at the old place drains and the new place continues, and a migration or a reshape does not need drain, destroy, and init. The only measured number behind that judgment, at the start, was an earlier result on A40s across InfiniBand: NCCL init plus the first collective took 493 to 523 ms, an order of magnitude above Llumnix’s migration pause. A back-of-the-envelope estimate said that a full rebuild is 55% to 57% of one reshape from TP1 to TP2 on an 8B model. Amortized over a switch every 2 s, that is 26%. Above 60 s it falls under 1%.

The dangerous baseline was already known. The plan said that if prebuilding every shape costs less than 5% of the KV budget and collectives do not regress, this mechanism is covered. The first step was not a protocol. It was to measure each fixed cost in the lifecycle.

What we measured

One process per GPU. A TCP star, independent of NCCL, is the barrier. Each timed call starts when the common barrier releases, and we report the maximum across ranks, so the number is the group’s critical path and not one rank’s local return. One unrecorded warmup, then ten recorded runs, median and p90. The clock is host wall time around the API call. Collectives include cudaStreamSynchronize. Every path that produces a new communicator also records the first 1 MB allreduce on that communicator.

The operations are init, split in half, shrink by dropping the last rank, grow by adding it back, finalize plus destroy, abort, reinit after abort, revoke, shrink with SHRINK_ABORT after revoke, buffer and window registration from 64 MB to 1 GB, a drain of four in-flight operations from 16 to 256 MB, and the matching c10d calls. The pool experiment on 4 GPUs builds a TP4 parent, two TP2 groups, and four TP1 groups. We measure the cudaMemGetInfo delta, a 64 MB allreduce on a child, latency with parent and child interleaved, and the time to capture one allreduce in a CUDA Graph.

Three machine groups. One node of four A40s, 48 GB, PCIe, pairs sharing a PCIe switch and the cross pair going over the CPU link. Two nodes of two A40s each, over InfiniBand, with no GPUDirect RDMA between GPU and NIC. One node of four V100 PCIe 32 GB, and two nodes of four V100s over InfiniBand. Software is NCCL 2.32.3 from the CUDA 12.9 wheel, the image’s 2.28.9 on A40 with torch 2.11, and 2.20.5 on V100 with torch 2.3. V100 stays on the CUDA 12 line because CUDA 13 no longer supports Volta. The anchor is a request-migration pause, median 69.1 ms.

On-demand rebuild puts a 50 to 180 ms lazy connect on the decode path. A prewarmed pool swaps the handle at a step boundary.
Figure 1. On-demand rebuild puts 50 to 180 ms of connect on the switch. A prewarmed pool pays that at startup. The switch is a handle swap and a graph recapture, and destroy of the old communicator runs in the background.

Figure 1 is the two switch paths. The top row is on-demand rebuild. Moving from TP4 to TP2 puts creation of the new group, plus its first collective, on the decode stream. The bottom row is a prewarmed pool. The TP2 communicator already exists and has already run one collective. The switch happens at a step boundary: swap the handle and recapture one CUDA Graph. Destroy of the old communicator goes to the background. The numbers in the figure come from the measurements below.

The money is in the first collective

Start with the cost of a new group. On four A40s with NCCL 2.32.3, init itself is 43 ms. Adding the first 1 MB allreduce makes it 92 ms. More than half is the first collective. Split is clearer. The split call is 23 ms. The first allreduce on the new subgroup pays another 54.5 ms. A later 64 MB allreduce on the same subgroup is 3.3 ms. NCCL postpones P2P and InfiniBand connect until first use, so however the new group was born, the first use connects.

Critical path of fresh init, split, shrink, grow, and abort-then-reinit, each plus the first collective, on A40 and V100.
Figure 2. On A40 and V100, on one node and across nodes, every path that creates a communicator plus its first collective lands between 43 and 195 ms, the same order as a 69 ms migration pause. A swap inside the pool is 0.5 ms.

Each group of bars is one path that produces a new communicator. Height is the median critical path including the first collective. The whisker reaches p90. The three fills are four A40s on one node, two nodes of two A40s, and two nodes of four V100s. The dashed line is the migration pause. The dotted line is a swap inside the pool. The important observation is how little split, shrink, and grow save against a fresh init. On four A40s, split is 81 ms, shrink 64 ms, grow 78 ms, and a fresh init 92 ms. On eight V100s, shrink is 180 ms, grow 175 ms, and fresh 184 ms. The only clear saving is a cross-node split, 43 ms, because cutting in half leaves each subgroup on one node and InfiniBand is no longer needed. splitShare=1 is 78 ms on four GPUs, against 81 ms without sharing. Sharing the parent’s resources does not save the first connect. All of these sit between 0.6x and 2.8x the migration pause. On-demand rebuild really is first order, which supports the premise of the conjecture.

Now the cost of tearing down. Abort and revoke sit at about 0.5 s in every configuration. Finalize plus destroy is 0.16 to 0.43 s.

Abort stays near 500 ms across NCCL versions, GPU types, and node counts. Finalize plus destroy does not grow with world size.
Figure 3. Abort stays between 485 and 534 ms on NCCL 2.28.9 and 2.32.3, on A40 and V100, on one node and across nodes. It does not move with world size. Destroy gets shorter as the world grows.

Each pair is abort and finalize-plus-destroy on one configuration. The dashed line is 500 ms. Abort barely moves from 2 GPUs to 8, from A40 to V100, or from 2.28.9 to 2.32.3. Revoke on A40 is also 499 to 502 ms. A floor that ignores scale, hardware, and version looks more like an internal wait than like real work. We did not confirm that in the source, so this is an inference. Destroy is shortest on eight V100s across nodes, 156 ms, which is the same point: it is not priced linearly in the number of connections. c10d’s destroy_process_group is more expensive, a stable 0.94 to 0.96 s on A40. This also corrects the original anchor. The 493 to 523 ms figure was a cold process’s first init plus first collective, including the first CUDA and library initialization inside the process. A warm process on the same class of cross-node A40 is 98 ms. The earlier number was high by about 5x.

Last, the pool. On four A40s, a TP4 parent plus two TP2 groups plus four TP1 groups costs an extra 464 MB per GPU, about 1% of memory. One TP2 child alone is about 76 MB. A 64 MB allreduce on the child is 3.273 ms, against 3.272 ms on a freshly built communicator. Five interleaved 16 MB collectives on parent and child are 13.35 ms, the same with and without sharing. The critical path of a switch inside the pool is one steady collective, about 0.2 ms, plus capturing and instantiating one allreduce in a CUDA Graph, 0.28 ms. That is about 0.7% of the migration pause.

Why the conjecture does not hold

The last two numbers from the pool decide it. The 50 to 180 ms of an on-demand rebuild is real, and it is large because the first connect was postponed until the switch. A prewarmed pool pays that same connect at startup, at 1% of memory, with no loss of bandwidth. Destroy and abort take 0.2 to 0.5 s and do not have to be waited on. They can finish in the background. Inference collectives close at the decode step, so the queue at a step boundary is already empty. Draining four large in-flight messages takes 7 to 112 ms on four A40s, but that cost does not sit on a step boundary.

The anchor itself is weaker than it looked. The cross-GPU request migration that motivated this is a handoff of state between two independent single-GPU engines. There is no communicator on that critical path. What remains as a real demand is two cases. The shape cannot be enumerated, for example moving a request onto a GPU that was not known in advance and forming a group with it, where grow’s 78 to 175 ms cannot be prepaid. Or in-flight messages cannot wait for a step boundary. This round did not find a real deployment of either case.

A few measurements are missing, and their direction matters. We measured PCIe A40s and V100s, not NVLink or NVSwitch. On NVLink the connect is faster and the pool is cheaper, which favors the conclusion. The pool was priced only at 4 GPUs. At 8 GPUs the shape count grows to about 15 and the memory to about 1 GB, still a few percent. Cross-node window registration crashes on 2.32.3 and was not measured. Buffer registration hits a cache, so the roughly 0 ms we read is not a real cost. Init on V100 with 2.20.5 takes 1.7 to 2.1 s, and we did not find out why. None of these sit on the switch, so they do not move the judgment.

Where the rebuild cost actually sits

Creating a new NCCL communicator spends its time on the first collective of that group, not on the init or split call. On four A40s, init is 43 ms, and 92 ms once the first allreduce is included. Split is 23 ms, and the first allreduce pays another 54.5 ms. That is PCIe A40 and V100, NCCL 2.28.9 and 2.32.3, with the default lazy connect. Any measurement of a rebuild has to include the first collective, or it undercounts by more than half.

Split, shrink, and grow are not much cheaper than a fresh init. At the same node count they land between 70% and 100% of a fresh init. splitShare=1 does not help the first connect. The exception is a split whose children no longer cross a node, which saves the InfiniBand connect. The machines and versions are the same as above. We did not measure NVLink or GPUDirect RDMA.

Abort and revoke sit on a floor of about 0.5 s from 2 to 8 GPUs, on A40 and V100, and on NCCL 2.28.9 and 2.32.3. Destroy is 0.16 to 0.43 s. c10d destroy is about 0.95 s. A switch can leave them in the background. Failure recovery has to close the old group first, and on that path the floor is a real cost.

When the shapes can be listed in advance, a prewarmed communicator pool is cheap. On 4 GPUs, TP4, TP2, and TP1 together cost an extra 464 MB per GPU, about 1% of 48 GB. Child collectives do not lose bandwidth. A switch inside the pool is sub-millisecond. This is one node.

A one-time lifecycle cost should be split into what can be paid before the event and what can be paid after it, and only then turned into a fraction of the event. Warmup, a prebuild, and a spare are paid before. Destroy in the background is paid after. Neither sits on the critical path. Only the part that can be neither prepaid nor deferred is the pain. Computed the original way, one reshape was 55% to 57% of the step. After the split it is 0.7%. That requires the prepaid resource to fit the budget, and the prepaid state not to expire before the event. If either fails, the split has to be done again.

A prewarmed pool already pays for the connections

Prebuilding a communicator for every shape on 4 GPUs, TP4, TP2, and TP1, costs an extra 464 MB per GPU, about 1% of memory. Collective bandwidth does not drop. The switch is sub-millisecond. The roughly 0.5 s floor of abort and revoke can run in the background. When the shapes can be enumerated, the primitives that already exist have already removed this cost. A rank-independent identity does not buy anything further.

RL Long Tail: Mid-Batch Parallel Shape Switching asks whether an RL rollout should leave DP4 for TP4 once only the long requests remain. On four PCIe A100s that switch costs at least 19 seconds, mostly a re-prefill, and a static DP2×TP2 finishes sooner.