Where Communicator Rebuild Actually Costs
Rebuilding an NCCL communicator on demand takes 50 to 180 ms, and almost all of that is the lazy connect inside the first collective. Warming a pool that uses about 1% of GPU memory moves that cost off the critical path.
Parallel shapes are starting to change at runtime
Large-model inference and training used to assume one thing. From the start of a distributed run to the end, the set of GPUs does not change. The degree of tensor parallelism, which splits one layer’s matrix by columns or by rows and merges with an allreduce at the end of the layer, is fixed at launch and then left alone. That assumption is loosening. Seesaw switches the parallel strategy between prefill and decode, because the two phases have very different compute-to-memory ratios and one degree cannot be best for both (Su et al., 2025). LoongServe lets the degree of sequence parallelism grow and shrink with request length (Wu et al., 2024). Llumnix migrates requests between instances at runtime and reports a pause of only tens of milliseconds (Sun et al., 2024). On the training side, torchft rebuilds the communicator from each step’s quorum when membership changes (PyTorch, torchft).
Once the set of participants can change, the communication layer becomes the problem. A collective library organizes communication around a communicator. A group of processes agrees on a unique id, each joins under a rank, and the library builds ring and tree topologies, allocates channel buffers, and opens intra-node P2P and inter-node InfiniBand connections. Rank is both the communication identity and the physical place. When membership changes, the old communicator no longer describes reality. The standard response today is to drain in-flight work, destroy the old group, and init a new one.
A few scales are worth holding onto. A 1 MB allreduce on four A40s is under a millisecond. 64 MB is about 3 to 7 ms. A decode step is usually tens of milliseconds. In an earlier experiment, a cross-GPU request migration paused for a median of 69 ms and happened about once every 2.6 s. If rebuilding a communicator takes hundreds of milliseconds, it is a first-order term in a single event.
What NCCL already offers for a changing group
Over the last few releases NCCL has been turning the communicator from a one-shot object into a mutable one. ncclCommSplit cuts a subgroup out of an existing communicator by color. With splitShare, the child can share the parent’s transport resources, which is meant to avoid reconnecting (NVIDIA, NCCL API). ncclCommShrink drops ranks from an existing communicator. With NCCL_SHRINK_ABORT it first stops in-flight work and then shrinks. ncclCommRevoke stops every in-flight operation and waits until the communicator is quiet, after which it can be destroyed, split, or shrunk. Since 2.29, ncclCommGrow adds ranks to an existing communicator, and together with shrink it follows nodes as they fail and return (NVIDIA, NCCL 2.29.2 release notes). ncclCommWindowRegister registers a buffer on a communicator for symmetric-memory kernels and the device API. The window is bound to that communicator, so a new communicator means a new registration.
These primitives avoid a full teardown when membership changes. They share a boundary. None of them keeps ordered in-flight messages. Revoke and SHRINK_ABORT drop work in progress. After a shrink, ranks are renumbered. The new communicator is a new object, and operations already queued on the old one do not move. PyTorch c10d wraps this with split_group, shrink_group, and new_group, and by default it creates the NCCL communicator lazily, on the first collective (PyTorch, torch.distributed).
There is also an old approach that is not an API. If the target shapes can be enumerated in advance, build a communicator for each of them at startup and switch handles. On a single node of 4 to 8 GPUs there are only a few reachable TP shapes. That path has always existed on paper. Nobody had priced it.
TL;DR
Inference and training are starting to change the parallel shape, and the set of GPUs, while a job is running. NCCL binds a communicator to a rank, so a new shape means tearing the old group down and building a new one. A natural guess is to detach identity from rank, and then a migration or a TP reshape would not need that teardown.
On A40 and V100, on one node and across InfiniBand, NCCL 2.20 through 2.32, an on-demand rebuild takes 50 to 180 ms. A request migration pauses for 69 ms, so the two are the same order. Almost all of the rebuild is the lazy connect inside the first collective. Split, shrink, and grow are not much cheaper than a fresh init. That connect can be paid at startup.
Prebuilding TP4, TP2, and TP1 on 4 GPUs costs an extra 464 MB per GPU, about 1% of memory. Collective bandwidth does not drop, and the switch is sub-millisecond. Abort and revoke sit on a floor of about 0.5 s, and that can run in the background. When the shapes can be listed in advance, the primitives that already exist have already removed this cost. A rank-independent identity does not buy anything further.
The conjecture we can test
The gap we thought we saw was the lack of a layer that separates a logical communication identity from a physical rank. If a logical flow has a stable endpoint, in-flight work at the old place drains and the new place continues, and a migration or a reshape does not need drain, destroy, and init. The only measured number behind that judgment, at the start, was an earlier result on A40s across InfiniBand: NCCL init plus the first collective took 493 to 523 ms, an order of magnitude above Llumnix’s migration pause. A back-of-the-envelope estimate said that a full rebuild is 55% to 57% of one reshape from TP1 to TP2 on an 8B model. Amortized over a switch every 2 s, that is 26%. Above 60 s it falls under 1%.
The dangerous baseline was already known. The plan said that if prebuilding every shape costs less than 5% of the KV budget and collectives do not regress, this mechanism is covered. The first step was not a protocol. It was to measure each fixed cost in the lifecycle.
What we measured
One process per GPU. A TCP star, independent of NCCL, is the barrier. Each timed call starts when the common barrier releases, and we report the maximum across ranks, so the number is the group’s critical path and not one rank’s local return. One unrecorded warmup, then ten recorded runs, median and p90. The clock is host wall time around the API call. Collectives include cudaStreamSynchronize. Every path that produces a new communicator also records the first 1 MB allreduce on that communicator.
The operations are init, split in half, shrink by dropping the last rank, grow by adding it back, finalize plus destroy, abort, reinit after abort, revoke, shrink with SHRINK_ABORT after revoke, buffer and window registration from 64 MB to 1 GB, a drain of four in-flight operations from 16 to 256 MB, and the matching c10d calls. The pool experiment on 4 GPUs builds a TP4 parent, two TP2 groups, and four TP1 groups. We measure the cudaMemGetInfo delta, a 64 MB allreduce on a child, latency with parent and child interleaved, and the time to capture one allreduce in a CUDA Graph.
Three machine groups. One node of four A40s, 48 GB, PCIe, pairs sharing a PCIe switch and the cross pair going over the CPU link. Two nodes of two A40s each, over InfiniBand, with no GPUDirect RDMA between GPU and NIC. One node of four V100 PCIe 32 GB, and two nodes of four V100s over InfiniBand. Software is NCCL 2.32.3 from the CUDA 12.9 wheel, the image’s 2.28.9 on A40 with torch 2.11, and 2.20.5 on V100 with torch 2.3. V100 stays on the CUDA 12 line because CUDA 13 no longer supports Volta. The anchor is a request-migration pause, median 69.1 ms.
Figure 1 is the two switch paths. The top row is on-demand rebuild. Moving from TP4 to TP2 puts creation of the new group, plus its first collective, on the decode stream. The bottom row is a prewarmed pool. The TP2 communicator already exists and has already run one collective. The switch happens at a step boundary: swap the handle and recapture one CUDA Graph. Destroy of the old communicator goes to the background. The numbers in the figure come from the measurements below.
The money is in the first collective
Start with the cost of a new group. On four A40s with NCCL 2.32.3, init itself is 43 ms. Adding the first 1 MB allreduce makes it 92 ms. More than half is the first collective. Split is clearer. The split call is 23 ms. The first allreduce on the new subgroup pays another 54.5 ms. A later 64 MB allreduce on the same subgroup is 3.3 ms. NCCL postpones P2P and InfiniBand connect until first use, so however the new group was born, the first use connects.
Each group of bars is one path that produces a new communicator. Height is the median critical path including the first collective. The whisker reaches p90. The three fills are four A40s on one node, two nodes of two A40s, and two nodes of four V100s. The dashed line is the migration pause. The dotted line is a swap inside the pool. The important observation is how little split, shrink, and grow save against a fresh init. On four A40s, split is 81 ms, shrink 64 ms, grow 78 ms, and a fresh init 92 ms. On eight V100s, shrink is 180 ms, grow 175 ms, and fresh 184 ms. The only clear saving is a cross-node split, 43 ms, because cutting in half leaves each subgroup on one node and InfiniBand is no longer needed. splitShare=1 is 78 ms on four GPUs, against 81 ms without sharing. Sharing the parent’s resources does not save the first connect. All of these sit between 0.6x and 2.8x the migration pause. On-demand rebuild really is first order, which supports the premise of the conjecture.
Now the cost of tearing down. Abort and revoke sit at about 0.5 s in every configuration. Finalize plus destroy is 0.16 to 0.43 s.
Each pair is abort and finalize-plus-destroy on one configuration. The dashed line is 500 ms. Abort barely moves from 2 GPUs to 8, from A40 to V100, or from 2.28.9 to 2.32.3. Revoke on A40 is also 499 to 502 ms. A floor that ignores scale, hardware, and version looks more like an internal wait than like real work. We did not confirm that in the source, so this is an inference. Destroy is shortest on eight V100s across nodes, 156 ms, which is the same point: it is not priced linearly in the number of connections. c10d’s destroy_process_group is more expensive, a stable 0.94 to 0.96 s on A40. This also corrects the original anchor. The 493 to 523 ms figure was a cold process’s first init plus first collective, including the first CUDA and library initialization inside the process. A warm process on the same class of cross-node A40 is 98 ms. The earlier number was high by about 5x.
Last, the pool. On four A40s, a TP4 parent plus two TP2 groups plus four TP1 groups costs an extra 464 MB per GPU, about 1% of memory. One TP2 child alone is about 76 MB. A 64 MB allreduce on the child is 3.273 ms, against 3.272 ms on a freshly built communicator. Five interleaved 16 MB collectives on parent and child are 13.35 ms, the same with and without sharing. The critical path of a switch inside the pool is one steady collective, about 0.2 ms, plus capturing and instantiating one allreduce in a CUDA Graph, 0.28 ms. That is about 0.7% of the migration pause.
Why the conjecture does not hold
The last two numbers from the pool decide it. The 50 to 180 ms of an on-demand rebuild is real, and it is large because the first connect was postponed until the switch. A prewarmed pool pays that same connect at startup, at 1% of memory, with no loss of bandwidth. Destroy and abort take 0.2 to 0.5 s and do not have to be waited on. They can finish in the background. Inference collectives close at the decode step, so the queue at a step boundary is already empty. Draining four large in-flight messages takes 7 to 112 ms on four A40s, but that cost does not sit on a step boundary.
The anchor itself is weaker than it looked. The cross-GPU request migration that motivated this is a handoff of state between two independent single-GPU engines. There is no communicator on that critical path. What remains as a real demand is two cases. The shape cannot be enumerated, for example moving a request onto a GPU that was not known in advance and forming a group with it, where grow’s 78 to 175 ms cannot be prepaid. Or in-flight messages cannot wait for a step boundary. This round did not find a real deployment of either case.
A few measurements are missing, and their direction matters. We measured PCIe A40s and V100s, not NVLink or NVSwitch. On NVLink the connect is faster and the pool is cheaper, which favors the conclusion. The pool was priced only at 4 GPUs. At 8 GPUs the shape count grows to about 15 and the memory to about 1 GB, still a few percent. Cross-node window registration crashes on 2.32.3 and was not measured. Buffer registration hits a cache, so the roughly 0 ms we read is not a real cost. Init on V100 with 2.20.5 takes 1.7 to 2.1 s, and we did not find out why. None of these sit on the switch, so they do not move the judgment.
Where the rebuild cost actually sits
Creating a new NCCL communicator spends its time on the first collective of that group, not on the init or split call. On four A40s, init is 43 ms, and 92 ms once the first allreduce is included. Split is 23 ms, and the first allreduce pays another 54.5 ms. That is PCIe A40 and V100, NCCL 2.28.9 and 2.32.3, with the default lazy connect. Any measurement of a rebuild has to include the first collective, or it undercounts by more than half.
Split, shrink, and grow are not much cheaper than a fresh init. At the same node count they land between 70% and 100% of a fresh init. splitShare=1 does not help the first connect. The exception is a split whose children no longer cross a node, which saves the InfiniBand connect. The machines and versions are the same as above. We did not measure NVLink or GPUDirect RDMA.
Abort and revoke sit on a floor of about 0.5 s from 2 to 8 GPUs, on A40 and V100, and on NCCL 2.28.9 and 2.32.3. Destroy is 0.16 to 0.43 s. c10d destroy is about 0.95 s. A switch can leave them in the background. Failure recovery has to close the old group first, and on that path the floor is a real cost.
When the shapes can be listed in advance, a prewarmed communicator pool is cheap. On 4 GPUs, TP4, TP2, and TP1 together cost an extra 464 MB per GPU, about 1% of 48 GB. Child collectives do not lose bandwidth. A switch inside the pool is sub-millisecond. This is one node.
A one-time lifecycle cost should be split into what can be paid before the event and what can be paid after it, and only then turned into a fraction of the event. Warmup, a prebuild, and a spare are paid before. Destroy in the background is paid after. Neither sits on the critical path. Only the part that can be neither prepaid nor deferred is the pain. Computed the original way, one reshape was 55% to 57% of the step. After the split it is 0.7%. That requires the prepaid resource to fit the budget, and the prepaid state not to expire before the event. If either fails, the split has to be done again.
A prewarmed pool already pays for the connections
Prebuilding a communicator for every shape on 4 GPUs, TP4, TP2, and TP1, costs an extra 464 MB per GPU, about 1% of memory. Collective bandwidth does not drop. The switch is sub-millisecond. The roughly 0.5 s floor of abort and revoke can run in the background. When the shapes can be enumerated, the primitives that already exist have already removed this cost. A rank-independent identity does not buy anything further.
Related analysis
RL Long Tail: Mid-Batch Parallel Shape Switching asks whether an RL rollout should leave DP4 for TP4 once only the long requests remain. On four PCIe A100s that switch costs at least 19 seconds, mostly a re-prefill, and a static DP2×TP2 finishes sooner.
按需重建一个 NCCL 通信组要 50 到 180 ms,但其中几乎全是首个 collective 的惰性建连,预热一个 1% 显存的多形状池就能把它挪出关键路径。
并行形状开始在运行中变化
大模型推理与训练过去默认一件事:一次分布式执行从开始到结束,参与的 GPU 集合不变。张量并行(tensor parallelism,TP,把一层的矩阵按列或行切到多张卡上,每层结束用 all-reduce 合并)的度数在启动时定下,之后不动。这个假设正在松动。Seesaw 在 prefill 与 decode 两个阶段之间切换并行策略,因为两个阶段的计算与访存比差得很远,同一套并行度不可能对两者都最优(Su et al. 2025)。LoongServe 让序列并行的度数随请求长度弹性伸缩(Wu et al. 2024)。Llumnix 在实例之间实时迁移请求,报告的迁移停机只有几十毫秒(Sun et al. 2024)。训练侧,torchft 用每个 step 的 quorum 在成员变化时重建通信组(PyTorch, torchft)。
一旦参与者集合会变,通信层就成了问题。集合通信库以 communicator 为单位组织通信:一组进程协商一个唯一 id,各自按 rank 编号加入,库在这组进程之间建立 ring 与 tree 拓扑、分配 channel 缓冲、建立节点内 P2P 与节点间 InfiniBand(IB)连接。rank 同时是通信身份与物理位置,成员一变,旧 communicator 就不再描述现实,今天的标准做法是 drain 掉在途操作、destroy 旧组、init 新组。
量级参照先放在这里。一次 1 MB 的 all-reduce 在 4 张 A40 上是亚毫秒,64 MB 约 3 到 7 ms;推理的 decode step 通常是几十毫秒;一次跨 GPU 的请求迁移在我们此前的实验里停顿中位 69 ms,平均每 2.6 s 发生一次。如果重建一个通信组要几百毫秒,它在单次事件里就是一阶项。
NCCL 已经给了哪些可变原语
NCCL 这几年一直在把 communicator 从一次性对象变成可变对象。ncclCommSplit 从一个已有 comm 按颜色切出子组,可以通过 splitShare 配置让子组共享父组的传输资源,意图是省掉重新建连(NVIDIA, NCCL API)。ncclCommShrink 从已有 comm 去掉若干 rank,带 NCCL_SHRINK_ABORT 标志时先终止进行中的操作再收缩;ncclCommRevoke 停掉所有在途操作并等 comm 静默,之后才能 destroy、split 或 shrink;2.29 起加入 ncclCommGrow,向已有 comm 动态加入 rank,与 Shrink 配合随节点失败与恢复调整成员(NVIDIA, NCCL 2.29.2 release notes)。ncclCommWindowRegister 把缓冲注册到 comm 上,供对称内存 kernel 与 device API 使用,window 绑在具体 comm 上,换 comm 就要重新注册。
这些原语解决的是成员变更时不必全量 teardown 的问题,共同的边界是都不保留有序的在途消息:Revoke 与 SHRINK_ABORT 丢弃进行中的操作,Shrink 后 rank 重新编号,新 comm 是新对象,旧 comm 上已入队的操作不会迁过去。PyTorch 的 c10d 在其上包了 split_group、shrink_group 与 new_group,默认惰性地在第一个 collective 时才真正创建 NCCL comm(PyTorch, torch.distributed)。
还有一条不在 API 里的老办法:如果目标形状能事先枚举,启动时就把每个形状的 communicator 建好,切换时换句柄。单节点 4 到 8 卡上可达的 TP 形状只有寥寥几种,这条路在纸面上一直存在,没人说清它的代价。
TL;DR
推理和训练开始在运行中改并行形状,也改参与的 GPU。NCCL 把通信身份绑在 rank 和 communicator 上,换形状就要拆掉旧组、建新组。一个自然的猜想是让身份脱离 rank,这样迁移和 TP reshape 就不必 teardown。
在 A40 与 V100、单节点与跨节点 IB、NCCL 2.20 到 2.32 上,按需重建要 50 到 180 ms,和一次 69 ms 的请求迁移停顿同量级。这笔时间几乎全在新组第一个 collective 的惰性建连上。split、shrink、grow 都不比全新 init 便宜多少。这步建连可以在启动时付掉。
为 4 卡上的 TP4、TP2、TP1 全部预建 communicator,每卡多占 464 MB,大约 1% 显存。collective 带宽不掉,切换是亚毫秒。abort 和 revoke 大约 0.5 s 的平台可以放到后台。形状能够事先列出来时,现成原语已经把这项成本消掉了。再做一层与 rank 无关的身份,并没有再多出一个机制。
可检验的猜想
我们看到的空白是:没有一层把逻辑通信身份与物理 rank 分开。如果一条逻辑通信流有稳定的 endpoint,切换时旧位置上的在途通信收口、新位置上继续,迁移与 reshape 就不必 drain、destroy、init。一开始支撑这个判断的唯一实测数字,是我们此前在 A40 跨节点 IB 上测到的 NCCL init 加首次 collective 493 到 523 ms,比 Llumnix 的迁移停机大一个数量级。量级估算显示:8B 模型 TP1 到 TP2 的一次 reshape 里,全量重建占 55% 到 57%;按 2 s 一次的切换间隔摊销是 26%,60 s 以上低于 1%。
当时已经知道最危险的对手是预建池。如果预建全部形状的显存低于 KV 预算 5% 且 collective 不退化,这个机制就被覆盖。所以第一步不是设计协议,而是把生命周期的每一项固定成本量出来。
我们测了什么
每 GPU 一个进程,用一个独立于 NCCL 的 TCP 星形控制面做 barrier,每个被计时的调用从共同 barrier 放行开始,报告各 rank 的最大值,这样测到的是整组的关键路径而不是某个 rank 的本地返回时间。每项 1 次不记录的预热加 10 次记录,报中位数与 p90。计时边界是 API 调用的 host 墙钟,collective 包含 cudaStreamSynchronize;凡是生成新 comm 的路径,都另记在新 comm 上的第一个 1 MB all-reduce。
覆盖的操作是 init、split(对半)、shrink(去掉最后一个 rank)、grow(把它加回来)、finalize 加 destroy、abort、abort 后 reinit、revoke、revoke 后 SHRINK_ABORT、缓冲注册与 window 注册(64 MB 到 1 GB)、4 条在途 16 到 256 MB 操作的 drain,以及 c10d 层的对应操作。池的实验在 4 卡上建 TP4 父 comm 加两个 TP2 加四个 TP1,量 cudaMemGetInfo 的增量、子 comm 上 64 MB all-reduce 与父子交错的延迟,以及 CUDA Graph 捕获一次 all-reduce 的时间。
硬件是三组:单节点 4 张 A40(48 GB,PCIe,两两同一 PCIe 交换,跨对走 CPU 间链路);两节点各 2 张 A40,经 IB 互联,GPU 与网卡之间没有 GPUDirect RDMA;单节点 4 张 V100 PCIe 32 GB 与两节点各 4 张 V100 经 IB。软件是 NCCL 2.32.3(CUDA 12.9 wheel)、镜像自带的 2.28.9(A40,torch 2.11)与 2.20.5(V100,torch 2.3)。V100 只能用 CUDA 12 这一路,因为 CUDA 13 已经不支持 Volta。对照锚点是一次请求迁移停顿,中位 69.1 ms。
这张示意图对比两条切换路径。上面一行是按需重建:从 TP4 切到 TP2 时,新组的创建加上它的第一个 collective 整段落在 decode 流上。下面一行是预热池:TP2 的 comm 在启动时已经建好并跑过一次 collective,切换发生在 step 边界,只需换句柄并重捕获一次 CUDA Graph,旧 comm 的 destroy 交给后台。图中的数字来自下一节的测量。
证据:钱花在第一个 collective 上
先看新组的代价。A40 4 卡上 NCCL 2.32.3 的 init 本身 43 ms,加上第一个 1 MB all-reduce 变成 92 ms,一半以上花在首次 collective 上。把 split 拆开看更清楚:split 调用本身 23 ms,新子组上第一个 all-reduce 另付 54.5 ms,而同一个子组之后的 64 MB all-reduce 只要 3.3 ms。NCCL 把 P2P 与 IB 的连接建立推迟到首次使用,所以不管新组是怎么来的,第一次用都要建连。
图中每组柱是一种生成新 comm 的路径,柱高是加上首个 collective 的关键路径中位数,误差线到 p90,三种填充分别是 A40 单节点 4 卡、A40 两节点 2x2 与 V100 两节点 4x2,虚线是迁移停顿,点线是池内切换。关键观察是 split、shrink、grow 相对全新 init 的差距很小:A40 4 卡上 split 81 ms、shrink 64 ms、grow 78 ms,fresh init 92 ms;V100 8 卡上 shrink 180 ms、grow 175 ms,fresh 184 ms。唯一明显便宜的是跨节点 split(43 ms),因为对半切之后每个子组都落在一个节点内,不再需要 IB 连接。splitShare=1 在 4 卡上是 78 ms,与不共享的 81 ms 没有区别,共享父组资源并不省首次建连。这些数都在迁移停顿的 0.6 到 2.8 倍,按需重建确实是一阶项,这一点支持猜想的前提。
再看拆的代价。abort 与 revoke 在所有配置上都停在约 0.5 s,finalize 加 destroy 是 0.16 到 0.43 s。
每组两根柱是同一配置下 abort 与 finalize 加 destroy 的中位数,虚线标 500 ms。abort 从 2 卡到 8 卡、从 A40 到 V100、从 2.28.9 到 2.32.3 几乎不变,revoke 在 A40 上也是 499 到 502 ms。一个与规模、硬件、版本都无关的平台,更像库内部某个等待超时而不是真实工作量,我们没有读源码确认,这里只作推断。destroy 在 V100 8 卡跨节点上反而最短(156 ms),同样说明它不是按连接数线性计价。c10d 层的 destroy_process_group 更贵,A40 上稳定在 0.94 到 0.96 s。这回过头也更正了最初的锚点:493 到 523 ms 是冷进程第一次 init 加首次 collective,含进程内首次 CUDA 与库初始化,而热进程在同类跨节点 A40 上是 98 ms,量级高估了约 5 倍。
最后看池。A40 4 卡上 TP4 父组加两个 TP2 加四个 TP1 共多占 464 MB/卡,约为显存的 1%;每个 TP2 子组单独约 76 MB。子组上 64 MB all-reduce 3.273 ms,与全新建的 comm 3.272 ms 相同;父子组交错跑 5 次 16 MB 是 13.35 ms,共享与不共享无差别。池内切换的关键路径是一次稳态 collective 约 0.2 ms,加 CUDA Graph 捕获并实例化一次 all-reduce 0.28 ms,约为迁移停顿的 0.7%。
为什么这个技术判断不成立
决定性的是池的最后两个数。按需重建的 50 到 180 ms 是真的,但它之所以大,是因为首次建连被推迟到了切换的时候。预热池把同一次建连挪到启动时,代价是 1% 显存,带宽不掉。destroy 和 abort 的 0.2 到 0.5 s 不必干等,可以在后台做完。推理的 collective 以 decode step 收口,step 边界本来就是空的。4 条在途大消息的 drain 在 A40 4 卡上要 7 到 112 ms,但 step 结束时队列里已经没有它们。
作为动机的跨 GPU 请求迁移,是在两个独立的单卡引擎之间交接状态。那条关键路径上没有 communicator。还可能真的需要重建的,只剩两种情况。一种是形状没法事先枚举,比如把请求迁到一张事先不知道的 GPU 上再和它建组,这时 grow 的 78 到 175 ms 无法预付。另一种是在途消息等不到 step 边界。这两种的真实部署,这一轮都没找到。
还有几处没测到。机器是 PCIe 的 A40 和 V100,没有 NVLink 和 NVSwitch。NVLink 上建连更快,池更便宜,这个方向对结论有利。池只算了 4 卡。8 卡的形状大约 15 个,显存到 1 GB 量级,仍是几个百分点。跨节点的 window 注册在 2.32.3 上崩溃,没有测。缓冲注册走缓存,读到的大约 0 ms 不是真实成本。V100 上 2.20.5 的 init 要 1.7 到 2.1 s,原因没查。这些都不在切换的关键路径上,不改变判断。
通信组重建真正贵在哪里
NCCL 生成新通信组,时间主要花在新组的第一个 collective 上,不在 init 或 split 调用本身。A40 4 卡上,init 是 43 ms,加上首次 all-reduce 变成 92 ms。split 是 23 ms,首次 all-reduce 再付 54.5 ms。这是在 PCIe 互联的 A40 与 V100、NCCL 2.28.9 与 2.32.3、默认惰性建连上测的。量任何重建方案,都要把首次 collective 算进去,否则会少计一半以上。
split、shrink、grow 并不比全新 init 便宜多少。节点数相同时,它们落在 fresh init 的 70% 到 100%。splitShare=1 也省不下首次建连。唯一明显更便宜的,是切完之后子组不再跨节点,这时省掉的是 IB 建连。机器和版本与上一条相同。NVLink 和 GPUDirect RDMA 没有测。
abort 与 revoke 在 2 到 8 卡、A40 与 V100、NCCL 2.28.9 与 2.32.3 上都停在大约 0.5 s。destroy 是 0.16 到 0.43 s,c10d 的 destroy 大约 0.95 s。切换时可以把它们放到后台,不必算进关键路径。故障恢复必须先把旧组收口,那条路径上这 0.5 s 就是真实成本。
形状能够事先列出来时,预热一个 communicator 池很便宜。4 卡上把 TP4、TP2、TP1 全部建好,每卡多占 464 MB,大约是 48 GB 显存的 1%。子组上的 collective 带宽不掉,池内切换是亚毫秒。这是单节点上测的。
一次性的生命周期成本,要先分开哪些能提前付、哪些能事后付,再算它占这次事件的比例。能在事件之前付掉的,例如预热、预建、备机,以及能在事件之后付掉的,例如后台 destroy,都不在关键路径上。只有两边都做不到的那一部分才是痛点。按原来的算法,单次占比是 55% 到 57%。分完之后只剩 0.7%。这要求预付的资源还在预算里,而且预付的状态在事件到来之前不会过期。这两条有一条不成立,就要重新分。
预热池已经把建连付掉了
最值得记住的是池。为 4 卡上 TP4、TP2、TP1 全部形状预建 communicator 只多占 464 MB/卡(约 1% 显存),collective 带宽不退化,切换只剩亚毫秒。abort 与 revoke 那约 0.5 s 的平台可以放到后台。形状能够事先枚举时,现成原语已经把这项成本消掉了。再做一层与 rank 无关的身份,并没有多出来的系统。
相关分析
RL 长尾上的并行形态切换 问的是 RL rollout 的长尾批,能不能在只剩长请求时从 DP4 切到 TP4。在四张 PCIe A100 上,一次切换至少 19 秒,大头是重新 prefill,静态的 DP2×TP2 更快。