Can Cross-Node Small-Message Communication Stay Resident?
About two thirds of the 13.6 us in an 8 B cross-node allreduce is control cost paid again on every operation. That cost can be paid once. An existing resident-proxy design already captures it, and two patches leave it only 1.6x from our upper bound.
During decode, a large model runs the whole network once per generated token. Tensor parallelism does one or two allreduces per layer. Expert parallelism does a dispatch and a combine, two all-to-alls per layer. Each message is a few KiB to a few tens of KiB. On a 100 Gb network the wire time for that size is a fraction of a microsecond, so almost all of the latency is fixed cost: kernel launch, proxy wakeup, buffer bounce, protocol handshake, and completion. Once the model is forced to span nodes, that fixed cost times dozens or hundreds of communications per step becomes a visible slice of decode latency. This post asks whether moving those costs from once per operation to once per epoch is a new system abstraction, and why the answer is no.
What one small cross-node message passes through
Start with the scale. Every measurement here is on NVIDIA EDR InfiniBand (100 Gb/s, ConnectX-5), with one GPU on each of two nodes. A CPU RDMA write ping-pong in host memory is about 1.15 us one way. That is the hardware floor of the wire, the switch, and the two NICs. Launching an empty kernel and synchronizing is about 6.4 us on one GPU. Replaying a single-node CUDA Graph and synchronizing is about 5.7 us. An 8 B device-to-host copy plus a sync is about 6 us. Every later number should be read against 1.15 us and these single-GPU costs.
NCCL is the de facto standard for GPU collectives (NVIDIA NCCL). Across nodes, the NCCL kernel on the GPU places data in a buffer and updates a counter. A host proxy thread polls that counter, and once it sees new work it posts an RDMA request to the NIC. The peer proxy sees the completion and notifies the peer kernel. Every communication walks this chain. CUDA Graphs capture a sequence of kernels and NCCL calls into one graph, so later steps replay the whole graph and the host launch cost becomes once per graph rather than once per operation (CUDA Graphs). Proxy wakeup, buffer handoff, and protocol synchronization are still paid every time.
GPUDirect RDMA (GDR) lets the NIC read and write GPU memory directly and skips the host bounce buffer (GPUDirect RDMA). The driver has to expose GPU pages to the NIC, either by loading nvidia_peermem into the RDMA peer-memory interface or by exporting DMA-BUF from the open kernel module and registering it. Without GDR, NCCL falls back to a host bounce buffer. The function is the same. The extra cost is a PCIe round trip.
Two lines of work already do this
Two lines are already moving fixed cost off the critical path.
The first is GPU-initiated networking. NVSHMEM organizes GPU memory into a symmetric partitioned global address space, and a kernel can call put and signal directly. Its IBGDA transport lets a GPU thread build the NIC work request and ring the doorbell, so the CPU is off the path (NVSHMEM). DeepEP builds a low-latency MoE all-to-all on top of it, and the receiver only waits on a signal (DeepEP). Since 2.28, NCCL also has a device API. GIN (GPU-Initiated Networking) lets a kernel post a network put and a signal directly. A symmetric window is registered once, and later operations only write the signal (NCCL Device API). On H100 with 400 Gb CX-7, published 8 B one-way numbers are 8.35 to 9.0 us for GIN and 8.0 to 12.15 us for NVSHMEM (Hamidouche et al., 2025). Those figures are half of a ping-pong round trip, and the paper does not say whether each round relaunches the kernel. The shared price of this line is GDR. A GPU can drive the NIC only if the NIC can reach GPU memory.
The second line is a resident host proxy. MSCCL++ splits communication into persistent channels such as PortChannel and MemoryChannel. The GPU writes a trigger into a host-visible FIFO. A spinning CPU proxy (ProxyService) reads it and posts an RDMA write and an atomic signal against pre-registered memory. The GPU spins on the semaphore (MSCCL++). The UCCL line also puts the transport on a resident CPU core, as a substitute for IBGDA (UCCL). What they remove is proxy wakeup and descriptor construction. What they keep is one GPU-to-host PCIe write and the NIC round trip. On the HPC side, the persistent collectives and partitioned communication in MPI 4.0 already express the same semantics: install the request once, then send one ready token per partition (MPI Forum, 2021).
Before any new code, the defining ability (install once, then only consume tokens) was already expressed by several existing primitives. The remaining question is empirical. On an ordinary cluster, do these primitives run, and how far do they get?
TL;DR
Cross-node collectives at decode size are dominated by a fixed cost paid on every call. The natural guess is to install the communication once, and after that only consume a ready token and a completion token.
On an EDR InfiniBand cluster without GPUDirect RDMA, a CUDA Graph NCCL allreduce of 8 B takes 13.6 us. The hardware floor is 1.15 us. About 8.9 us is NCCL’s per-operation control. A resident host proxy can reach 4.7 us, so the gain is real. MSCCL++ PortChannel has the same shape, but it will not start, because it treats pinned host memory as GPU memory. Two patches later it is 9.8 us on V100, 1.6 times our bound. The rest is fences, buffer classification, and polling. One NCCL switch, shared buffers off, also halves 8 B send/recv.
An end-to-end estimate from the measured curves saves more than 10% only for small-batch cross-node MoE expert-parallel decode. The resident proxy is already there. What remains is tuning.
The conjecture we can test
The starting point was a number from an earlier measurement. On one pair of A40 nodes, a CUDA Graph NCCL send/recv was 37.7 us one way, while an RDMA write in host memory was 1.13 us. The wire was 3% of the time. The rest sat in the endpoints. If a communication that repeats inside an epoch, with a fixed shape and a fixed peer, becomes a resident object, then queue-pair setup, memory registration, the proxy thread, and a persistent kernel are paid once at install time. After that, each use is one ready signal and one completion signal, and most of the fixed cost on the decode path can be amortized. The optimistic bound at the time was 10% to 18% of a decode step in the most favorable case.
The judgment then was that 10% to 18% is worth measuring no matter who it belongs to. Missing GDR might make part of the cost a per-message bounce that residency cannot amortize. Another part is protocol and proxy, which is what GIN and IBGDA aim at. The net gain of this mechanism over the strongest existing combination is the gap between a resident state machine and those existing paths on the same hardware. That gap had no data.
What we measured
There are two pairs of machines. One pair is two servers, each with one NVIDIA A40 (Ampere, 48 GB), EDR InfiniBand, NCCL 2.28.8, and CUDA 13.0. The other pair is two servers, each with one V100 (Volta, 32 GB), on the same EDR fabric. CUDA 13 no longer supports Volta, so that pair uses CUDA 12.4, NCCL 2.20.5, and UCX 1.16. On both pairs NCCL reports GDR disabled in every configuration, and forcing the GDR level does not change that. A survey of all 20 GPU nodes in the cluster found ConnectX-5 NICs everywhere. nvidia_peermem is installed on every node and loaded on none. The datacenter GPUs use the closed kernel module, so they have no DMA-BUF. The one card that does have DMA-BUF has no second node to pair with. Every conclusion here is limited to an environment without GDR.
We wrote a resident host-proxy upper bound, called B4 below. A persistent kernel writes a flag mapped into host memory. A spinning CPU thread sees it and posts one RDMA write to already-registered memory. The peer’s persistent kernel polls the arriving flag and data. This is the simplest form of the MSCCL++ PortChannel family. In a second round we added byte-for-byte checks on the receiver and acquire-semantics polling. Every iteration was zero errors.
The metric is the median one-way time, a ping-pong round trip divided by two, for messages from 8 B to 64 KiB. The GPU side uses globaltimer or CUDA events. Each point is 1000 iterations, with one warmup and then five recorded runs. The controls are NCCL allreduce and send/recv under Graph and eager, plus NCCL knobs for protocol, channel count, shared buffers, launch mode, proxy batching, and the network plugin. We also ran UCX put, tag, and active message on GPU memory, MSCCL++ PortChannel with ProxyService, NCCL GIN, NVSHMEM with ibrc, ibgda, and ucx remote transports, and the UCCL p2p engine. The measured latency curves were then dropped into several decode shapes for a closed-form end-to-end estimate.
Where the 13.6 us goes
The left side of Figure 1 is the two paths. The top row is NCCL after a CUDA Graph capture. Each operation still enters the kernel, waits for the proxy to notice new work, bounces through host memory when there is no GDR, posts the network request, synchronizes the protocol, and completes in the peer kernel. The bottom row is the resident proxy. The kernel and the proxy thread are already resident. Each operation is one flag write, one RDMA write, and one poll.
The right side breaks down the 13.6 us of an 8 B Graph allreduce on the A40 pair. The gray 1.15 us is the hardware floor. The blue 2.7 us is the PCIe round trip between GPU and host without GDR, from B4’s flag-only 3.76 us minus the floor. The green 0.9 us is the extra cost of moving a 16 B payload through mapped memory. The remaining hatched 8.85 us is NCCL’s own per-operation control: proxy wakeup, bounce, protocol sync, and kernel entry. B4’s 4.66 us sits at the dashed line. This supports the first-round judgment. About 66% of small-message latency is control cost that can move from once per operation to once per epoch, and a CUDA Graph does not get it, because a Graph only amortizes the host launch. Eager send/recv is more extreme. Of 27.6 us, launch plus sync is only 3.3 us. The other 19.6 us sits on kernel boundaries, the proxy, and the bounce.
Two surprises showed up here. The earlier 37.7 us depends on the node pair. The same image and the same method on another pair is 24.3 us, so a number carried across pieces of work has to be remeasured on the pair it will be compared against. Most NCCL knobs do nothing at 8 B. The LL protocol matches the default. LL128 and Simple are slower. Grouped launch, proxy batching, and a different network plugin do not move the number. At 64 KiB, B4 is 18.0 us and the NCCL Graph allreduce is 37.0 us, so the gap shrinks to about 2x. Fixed-cost dominance holds only below a few KiB.
At this point the judgment passed the first gate. The gain is real, and it is large. The second round asks whether an existing mechanism already takes that gain.
How much existing mechanisms already take
Figure 2 is every mechanism that ran on the V100 pair. The axis is the median one-way time at 8 to 16 B. The parentheses are the ratio to the validated B4. The corner lists implementations that did not initialize. Gray is the 1.1 us hardware floor. Green is two fence variants of B4. Blue is MSCCL++. Red is UCX and NCCL.
The GDR-dependent family cannot initialize in this environment. NCCL GIN reports neither peermem nor DMA-BUF, and both the device API and global GIN support are 0. NVSHMEM ibrc fails to build its transport table after the DMA-BUF check fails. ibgda reports that the peer GPU is not reachable. ucx reports Bad address while registering the GPU heap. The UCCL p2p engine only offers the two GPU-memory registration paths that need peermem or DMA-BUF, and it has no host-bounce mode. DeepEP sits on NVSHMEM and is unavailable with it. These failures point at GPU-memory registration, not at the communication mechanism, so we record them as unmeasurable here rather than as open space.
MSCCL++, which has the same shape as B4, also failed at first, with an error that nvidia_peermem was not loaded. The source shows that the design does not require GDR. It classifies a buffer by asking which device owns the address, and pinned host memory from cudaHostAlloc is classified as GPU memory, so registration takes the peermem branch. Its semaphore also lives in GPU memory, and the proxy’s RDMA atomic on that semaphore needs GDR too. Two changes fix it. Register host memory as host memory according to the CUDA pointer attributes, and put the semaphore in mapped pinned host memory. PortChannel plus ProxyService then runs, with zero errors on a byte-for-byte check, 9.8 us one way at 16 B, and 31.3 us at 64 KiB.
The important distance in Figure 2 is that row against B4: 6.1 versus 9.8 us, a factor of 1.59. B4 itself, rewritten with a per-thread system fence, is 7.8 us, and against that conservative version MSCCL++ is only 1.2x slower. Of the 3.6 us between them, a substantial part is the fence. The rest is buffer classification, proxy polling, and how the signal is posted, which are implementation differences inside one design. On the NCCL side, NCCL_NET_SHARED_BUFFERS=0 gives each peer its own bounce buffer. An 8 B send/recv then drops from the default 23.5 us to 11.6 us. One switch halves it, and it is the largest NCCL knob we found. No knob pushes allreduce below about 15 us. A UCX active message on GPU memory is 10.6 us.
How much of the saved time reaches decode
Figure 3 substitutes the measured A40 curves for Graph NCCL and for B4 into several decode shapes: Llama-3.1-8B and 70B with tensor parallelism across two nodes, and Qwen3-30B-A3B with 8-way expert parallelism across two nodes, at batch 1, 8, and 64. The model shape sets the number of communications and the message size per step. Compute time comes from a calibration on the same class of A40. The vertical axis is the fraction of the step saved by replacing NCCL with B4, and B4’s advantage is capped at 64 KiB, because we did not measure it still leading on larger messages. This is a closed-form estimate, not an end-to-end run. Expert-parallel all-to-all is approximated as one point-to-point times the number of peers, which biases the savings upward.
Only two kinds of cell cross the 10% line: Qwen3-30B-A3B expert parallelism saves 22.9% at batch 1 and 11.3% at batch 8, and Llama-3.1-8B tensor parallelism of 4 across nodes saves 12.4% at batch 8. The 70B models are compute-heavy, and the saved fraction stays under 4%. At batch 64 the message exceeds 64 KiB and the gain falls below 5%. Pipeline parallelism across nodes has only two communications per step, so the fraction is 0 and is not drawn. The baseline here is the B4 upper bound. Against MSCCL++, which already exists, or against NCCL with the knob turned, the per-operation room shrinks from about 9 us to about 3.6 us. A normal scheduler also tries not to put tensor parallelism across nodes. A search on the same class of A40 chooses tensor parallelism inside a node and data parallelism across nodes, so the tensor-parallel columns themselves sit away from a real deployment.
Why the conjecture does not hold
The first half holds. On an ordinary InfiniBand cluster without GDR, about two thirds of an 8 B cross-node communication is control paid on every operation, and that control can be paid once. A CUDA Graph does not get it. A resident proxy does. The second half does not hold. This is not a missing capability. MSCCL++ PortChannel already has the resident proxy, the trigger FIFO, and the atomic signal. It failed to run here because of the environment and one buffer-classification detail. After that detail is patched, it sits 1.2 to 1.6 times from our bound. The rest is fences, polling, and adaptation. That is a tuning pass, not a new abstraction. The end-to-end saving is also narrow. It shows up in small-batch cross-node expert parallelism.
A few measurements are missing. With GDR we never compared GIN or IBGDA with B4 on this hardware. The published numbers come from a faster network and a faster GPU, and the timing boundary is unclear, so we cannot say whether the device-initiated path has reached this bound. Those two lines were built for this problem, so the unmeasured part is more likely to shrink the remaining room than to enlarge it. The patched MSCCL++ payload lives in host memory, so every GPU read and write crosses PCIe. That makes the number slower, which favors our conclusion, and 1.6 times is an upper bound on the gap. The end-to-end estimate treats all-to-all as point-to-point and biases the savings upward. The V100 and A40 numbers are not on the same pair of nodes. Comparisons stay inside one pair. The first-round B4 had no byte-for-byte check. After the second round added one, the number moved from 6.19 to 6.14 us. The earlier 4.7 us is not an artifact. It is an A40 number, and it cannot be set next to MSCCL++ on V100.
What these measurements support directly
The bottleneck of small cross-node messages is endpoint software, and mostly the control steps paid again on every operation. On EDR InfiniBand without GDR, the hardware is 1.15 us of an 8 B Graph allreduce of 13.6 us. The PCIe round trip plus payload is 3.6 us. The remaining 8.9 us is proxy wakeup, bounce, and protocol sync. A CUDA Graph only removes the host launch, and it never touches this part. The claim holds below a few KiB. At 64 KiB the gap is already down to about 2x.
On a cluster without GDR, the newer device-initiated stacks fail as a group, and the error does not say that the cause is driver configuration. Ordinary NCCL silently falls back to a host bounce, so day-to-day training and inference look fine. GIN, NVSHMEM, UCCL, and DeepEP fail during initialization. Before arguing about the mechanism, check whether nvidia_peermem is loaded, whether the driver is the open kernel module, and whether the CUDA properties report GDR and DMA-BUF support.
An existing system that matches your design, and that only fails to run because of the environment, is still your baseline. MSCCL++ looked like an unsupported environment. It was a buffer-classification detail, and two edits brought it next to our bound. Stopping at the error would have turned a tuning problem into a missing capability. The other case is different. When the existing mechanism itself depends on the missing capability, as IBGDA does on GDR, its absence only means neither side was measured. It does not assign the space to a new design.
For small send/recv, try NCCL_NET_SHARED_BUFFERS=0 first. On our V100 pair it cuts 8 B one-way time from 23.5 us to 11.6 us, more than every other knob combined. Allreduce is insensitive to it. The spread across node pairs (37.7 versus 24.3 us under the same method) is also larger than most knobs, so a comparison of designs has to hold the node pair fixed.
The end-to-end value of residency depends on the ratio of communications per step to compute, not on how many times a single latency improves. Saving 2x to 3x per operation is 2% to 4% on a 70B tensor-parallel decode, and 11% to 23% on small-batch cross-node expert parallelism. Count the communications and the message sizes of the target parallel scheme first, then multiply by the measured per-operation gap.
The mechanism is already there
Only small-batch cross-node MoE expert-parallel decode saves more than 10%. The resident proxy is already there. What remains is tuning.
8 B 跨节点 allreduce 的 13.6 us 里有三分之二是每次操作重付的控制成本,这部分确实能变成一次性的,但现成的常驻 proxy 设计已经拿到了它,打两处补丁就只差 1.6 倍。
大模型推理的 decode 阶段每生成一个 token 都要跑一遍整个模型,张量并行下每层要做一两次 all-reduce,专家并行下每层要做 dispatch 与 combine 两次 all-to-all,每次搬的数据只有几 KiB 到几十 KiB。这种规模的消息在 100 Gb 网络上的传输时间是零点几微秒,于是一次通信的耗时几乎全是固定成本:内核启动、代理线程唤醒、缓冲区中转、协议握手与完成通知。模型一旦被迫跨节点切分,这些固定成本乘上每 step 几十到上百次通信,就成了 decode 延迟里看得见的一块。这篇讲我们如何检验把这些固定成本从每次付一次挪到每个 epoch 付一次是否构成一个新的系统抽象,以及为什么最后的答案是否定的。
一次跨节点小消息要经过什么
先交代量级。本文所有测量都在 NVIDIA EDR InfiniBand(100 Gb/s,ConnectX-5 网卡)上,两个节点各出一张 GPU。CPU 直接在主机内存上做一次 RDMA write 的 ping-pong,单程约 1.15 us,这是线路、交换机与两端网卡加起来的硬件地板。单卡上发射一个空 kernel 并同步约 6.4 us,把一个单节点 CUDA Graph 重放并同步约 5.7 us,一次 8 B 的设备到主机拷贝加同步约 6 us。后文所有数字都要跟 1.15 us 与这几个单卡数比。
NCCL 是 GPU 集合通信的事实标准(NVIDIA NCCL)。跨节点时,GPU 上的 NCCL kernel 负责把数据放进缓冲区并更新计数,主机上一个代理线程(proxy)轮询这些计数,发现有新工作后向网卡提交 RDMA 请求,对端代理收到完成事件后再通知对端 kernel。每次通信都要走一遍这条链。CUDA Graph 允许把一串 kernel 与 NCCL 调用捕获成一张图,之后每次只重放整张图,主机侧的逐个发射开销因此从每次操作变成每张图一次(CUDA Graphs),但代理唤醒、缓冲区交接与协议同步仍然每次都付。
GPUDirect RDMA(下称 GDR)让网卡直接读写 GPU 显存,省掉经主机内存的中转(GPUDirect RDMA)。它需要驱动层把显存页暴露给网卡,要么加载 nvidia_peermem 模块接入 RDMA 栈的 peer memory 接口,要么由开源内核模块导出 DMA-BUF,再用对应的注册接口。没有 GDR 时 NCCL 会退回到经主机 bounce buffer 中转,功能不变,只是多了 PCIe 往返。
已经有人在做同一件事
把固定成本挪出关键路径有两条正在并行推进的路线。
第一条是 GPU 发起网络。NVSHMEM 把多 GPU 的显存组织成一个对称的分区全局地址空间,kernel 里可以直接调用 put 与 signal;它的 IBGDA 传输让 GPU 线程自己构造网卡的工作请求并敲门铃,CPU 完全不在路径上(NVSHMEM)。DeepEP 在其上实现了面向 MoE 的低延迟 all-to-all,接收端只等信号(DeepEP)。NCCL 从 2.28 起也提供 device API,其中 GIN(GPU-Initiated Networking)让 kernel 内直接发网络 put 与 signal,对称窗口注册一次后每次只写信号(NCCL Device API)。在 H100 加 400 Gb CX-7 的系统上,公开报告的 8 B 单程是 GIN 8.35 到 9.0 us、NVSHMEM 8.0 到 12.15 us(Hamidouche et al. 2025),那组数字是 ping-pong 往返的一半,是否包含每轮重新发射 kernel 文中没有写明。这条路线的共同代价是离不开 GDR:GPU 直接驱动网卡的前提就是网卡能直接访问显存。
第二条是常驻 host proxy。MSCCL++ 把通信拆成 PortChannel、MemoryChannel 这类持久通道,GPU 往一个主机可见的 FIFO 写一个触发项,一个常驻自旋的 CPU 代理(ProxyService)读到后对预先注册的内存发 RDMA write 与原子信号,GPU 侧自旋等信号量(MSCCL++)。UCCL 系列同样把传输逻辑放在常驻 CPU 核上,作为 IBGDA 的替代(UCCL)。它们消除的是代理唤醒与描述符构建,保留的是一次 GPU 到主机的 PCIe 写与网卡往返。更早的 HPC 侧,MPI 4.0 的持久集合操作与分区通信早已在语义上表达了请求安装一次、每个分区只发一个就绪令牌(MPI Forum 2021)。
所以动手之前就知道,定义性的能力(安装一次、之后只消费令牌)已经被好几种现成原语表达了。剩下的问题是经验性的:在一个普通集群上,这些原语到底跑不跑得起来、跑到多少。
TL;DR
decode 规模的跨节点 collective,时间几乎全是每次调用都要重付的固定成本。自然的猜想是把通信装好一次,之后每次只消费一个就绪令牌和一个完成令牌。
在没有 GPUDirect RDMA 的 EDR IB 上,CUDA Graph 捕获的 8 B NCCL allreduce 是 13.6 us。硬件地板只有 1.15 us。大约 8.9 us 是 NCCL 每次操作的控制。常驻 host proxy 可以到 4.7 us,所以这笔收益是真的。同构的 MSCCL++ PortChannel 一开始跑不起来,因为它把 pinned 主机内存判成了显存。改两处之后,V100 上是 9.8 us,离上界 1.6 倍。剩下的是 fence、buffer 分类和轮询。NCCL 关掉共享 buffer,也能把 8 B send/recv 减半。
按实测曲线做的端到端估算里,只有小 batch 跨节点 MoE 专家并行的 decode 能省下 10% 以上。常驻 proxy 已经在。剩下的是调优。
可检验的猜想
我们的出发点是上一次测量留下的一个数:同一类 A40 两节点上,CUDA Graph 捕获的 NCCL send/recv 单程 37.7 us,而主机内存上的 RDMA 只要 1.13 us。线路只占 3%,其余全在端点。如果把一段在 epoch 内反复发生、形状与对端都不变的通信做成常驻对象,建立时一次性付掉队列对建立、内存注册、代理线程与 kernel 常驻,之后每次只剩一次就绪信号与一次完成信号,那么 decode 路径上的大部分固定成本就可以摊掉。当时估的乐观上界是最有利场景下 decode step 的 10% 到 18%。
当时的判断是,这 10% 到 18% 不管属于谁,都值得先测清楚。GDR 缺失可能让一部分成本是按消息付的中转,常驻化摊不掉;另一部分是协议与代理,正是 GIN 与 IBGDA 瞄准的部分。这个机制相对最强现成组合的净收益,等于常驻状态机与这些现成路径在同一硬件上的差,这个差当时没有数据。
我们测了什么
硬件有两对。一对是两台各出一张 NVIDIA A40(Ampere,48 GB)的服务器,EDR IB,NCCL 2.28.8,CUDA 13.0;另一对是两台各出一张 V100(Volta,32 GB)的服务器,同样的 EDR IB,因为 CUDA 13 已不支持 Volta,用 CUDA 12.4、NCCL 2.20.5 与 UCX 1.16。两对上 NCCL 在所有配置下都报告 GDR 未启用,强制 GDR 级别也不变。我们对集群的全部 20 个 GPU 节点做了普查:网卡都是 ConnectX-5,每台都装了 nvidia_peermem 模块但都没加载,datacenter GPU 用的是闭源内核模块因而没有 DMA-BUF,唯一有 DMA-BUF 的一张卡凑不出第二个节点。所以本文的结论全部限定在无 GDR 的环境。
我们自己写了一个常驻 host proxy 上界,下称 B4:一个持久 kernel 往映射到主机的 flag 写一次,一个常驻自旋的 CPU 线程看到后对已注册的内存发一次 RDMA write,对端的持久 kernel 轮询到达的 flag 与数据。它就是 MSCCL++ PortChannel 那一族设计的最简形式。第二轮我们给它补上了接收端逐字校验与 acquire 语义的轮询,所有迭代 0 错。
测量口径统一为 ping-pong 往返除以二的单程中位数,消息 8 B 到 64 KiB,GPU 侧用 globaltimer 或 CUDA event 计时,每点 1000 次迭代,预热 1 次后记 5 次。对照组包括 Graph 与 eager 下的 NCCL allreduce 与 send/recv,以及 NCCL 的协议、channel 数、共享 buffer、发射模式、代理批量与网络插件等旋钮;UCX 的 put、tag 与 active message 在 GPU 内存上的三种走法;MSCCL++ PortChannel 加 ProxyService;NCCL GIN;NVSHMEM 的 ibrc、ibgda 与 ucx 三种远端传输;UCCL 的 p2p 引擎。最后把实测的延迟曲线代进几种模型的 decode 形状,做一个端到端的闭式估算。
13.6 us 去了哪里
图 1 左侧是两条路径的示意。上面一行是 CUDA Graph 捕获后的 NCCL:每次操作仍要进入 kernel、让代理线程发现新工作、在无 GDR 时经主机缓冲区中转、发网络请求并做协议同步、最后让对端 kernel 完成。下面一行是常驻 proxy:kernel 与代理线程都已常驻,每次只有一次 flag 写、一次 RDMA write 与一次轮询。
右侧是 A40 对上 8 B Graph allreduce 13.6 us 的拆解。灰色 1.15 us 是硬件地板;蓝色 2.7 us 是无 GDR 时 GPU 与主机之间的 PCIe 往返,由 B4 只写 flag 的 3.76 us 减去地板得到;绿色 0.9 us 是经映射内存搬 16 B payload 的增量;剩下斜线的 8.85 us 归给 NCCL 自身的每次操作控制,包括代理唤醒、中转、协议同步与 kernel 入口。B4 的 4.66 us 在虚线处。这张图支持第一轮的判断:约 66% 的小消息延迟是可以从每次付挪到每个 epoch 付的控制成本,而 CUDA Graph 拿不到它,因为 Graph 只摊掉了主机发射。eager 模式的 send/recv 更极端,27.6 us 里发射加同步只有 3.3 us,其余 19.6 us 都在 kernel 边界、代理与中转上。
两个出乎预期的观察也在这一步。上次那个 37.7 us 与节点对有关,同镜像同口径换一对节点只有 24.3 us,所以跨工作引用这类数时必须带着节点对重测。NCCL 的大多数旋钮对 8 B 没有作用:LL 协议与默认相同,LL128 与 Simple 反而更慢,分组发射、代理批量与换网络插件都不改变数字。到 64 KiB 时 B4 是 18.0 us,NCCL Graph allreduce 是 37.0 us,差距缩到约 2 倍,固定成本主导只在 KiB 以下成立。
到这里,这个判断通过了第一道闸口:收益是真实的,而且大。第二轮要回答的是现成机制是否已经拿到这部分。
现成机制拿到了多少
图 2 是 V100 对上所有跑起来的机制,横轴是 8 到 16 B 单程中位,括号里是相对校验版 B4 的倍数,右上角列出无法初始化的实现。灰色是硬件地板 1.1 us,绿色是 B4 的两个 fence 变体,蓝色是 MSCCL++,红色是 UCX 与 NCCL。
依赖 GDR 的一族在这个环境里全部无法初始化。NCCL GIN 报告既不支持 peermem 也不支持 DMA-BUF,device API 与全局 GIN 支持都是 0;NVSHMEM 的 ibrc 在 DMA-BUF 检查失败后建传输表失败,ibgda 报对端 GPU 不可访问,ucx 在注册 GPU 堆时报 Bad address;UCCL 的 p2p 引擎只提供依赖 peermem 或 DMA-BUF 的两条显存注册路径,没有主机中转模式;DeepEP 建在 NVSHMEM 上,随之不可用。这些失败指向的都是显存注册这个前提本身,不是它们的通信机制,所以我们把它们记为本环境不可测,而不是把它们的缺席当作空间。
与 B4 同构的 MSCCL++ 最初也失败了,报错是 nvidia_peermem 未加载。读源码后发现,这并不是它的设计离不开 GDR:它用地址查询设备号来判断 buffer 类型,cudaHostAlloc 分配的 pinned 主机内存也会被判成显存,于是走了需要 peermem 的注册分支;另外它的信号量放在显存里,代理要对它做 RDMA 原子操作,同样需要 GDR。改两处,一是按 CUDA 指针属性把主机内存按主机内存注册,二是把信号量放进映射的 pinned 主机内存,PortChannel 加 ProxyService 就跑起来了,逐字校验 0 错,16 B 单程 9.8 us,64 KiB 31.3 us。
图 2 的关键观察是这一行与 B4 的距离:6.1 对 9.8 us,1.59 倍。而 B4 自己换成每线程 system fence 的保守写法就是 7.8 us,对这个保守版本 MSCCL++ 只慢 1.2 倍。也就是说两者之间 3.6 us 的差距里,相当一部分可以由 fence 的选择解释,其余在 buffer 分类、代理轮询与信号方式上,属于同一设计的实现差异。NCCL 这一侧,打开 NCCL_NET_SHARED_BUFFERS=0 让每个对端使用独占的中转 buffer 之后,8 B send/recv 从默认的 23.5 us 降到 11.6 us,一个开关就减半,是我们找到的效果最大的 NCCL 旋钮;allreduce 则没有任何旋钮能压到约 15 us 以下。UCX 的 active message 在 GPU 内存上是 10.6 us。
省下来的时间有多少能落到 decode 上
图 3 把 A40 上 Graph NCCL 与 B4 的实测延迟曲线代进几种 decode 形状:Llama-3.1-8B 与 70B 的张量并行跨两节点,以及 Qwen3-30B-A3B 的 8 路专家并行跨两节点,batch 为 1、8、64,每 step 的通信次数与消息大小由模型形状给出,计算时间取自同一类 A40 上的校准。纵轴是把 NCCL 换成 B4 后省下的 step 比例,并把 B4 的优势限制在 64 KiB 以内,因为更大的消息我们没有测到它仍然领先。这是闭式估算,不是端到端实测,专家并行的 all-to-all 用单次点对点乘次数近似,这个近似偏向高估收益。
只有两类格子越过 10% 虚线:Qwen3-30B-A3B 专家并行 batch 1 省 22.9%、batch 8 省 11.3%,以及 Llama-3.1-8B 张量并行 4 路跨节点 batch 8 的 12.4%。70B 模型的计算太重,省下的比例都在 4% 以内;batch 64 时消息超过 64 KiB,收益掉到 5% 以下;流水线并行跨节点每 step 只有两次通信,比例是 0,没画出来。还要注意,这里的对照是 B4 这个上界,换成已经存在的 MSCCL++ 或调过旋钮的 NCCL,每次操作可争取的余量从约 9 us 缩到约 3.6 us。另外,正常的调度会尽量避免让张量并行跨节点,同一类 A40 上的顺序搜索会选择节点内张量并行加节点间数据并行,所以张量并行这几列本身就偏离真实部署。
为什么这个判断不成立
前半句成立。在没有 GDR 的普通 IB 集群上,8 B 跨节点通信大约三分之二是每次操作都要重付的控制,这部分可以变成一次性成本。CUDA Graph 拿不到它。常驻 proxy 可以。后半句不成立。这不是一种缺了的能力。MSCCL++ PortChannel 已经有常驻代理、触发 FIFO 和原子信号。它在这个集群上跑不起来,是环境和一处 buffer 分类。补上之后,离上界 1.2 到 1.6 倍。剩下的是 fence、轮询和适配,是一次调优,不是一个新的抽象。端到端能省的时间也很窄,只出现在小 batch 跨节点专家并行里。
还有几处没测到。有 GDR 时,GIN 和 IBGDA 相对 B4 在这台机器上没有数。公开数字来自更快的网络和 GPU,计时边界也不清楚,所以不能说设备发起的路径已经到了这个上界。那两条路线本来就是为这件事设计的,没测到的部分更可能让剩余空间变小。打过补丁的 MSCCL++,payload 在主机内存里,GPU 每次读写都要过 PCIe,数字会偏慢。这个方向对结论有利,所以 1.6 倍是差距的上界。端到端估算把 all-to-all 当成点对点,会把收益估高。V100 和 A40 不在同一对节点上,比较只在同一对内部做。第一轮的 B4 没有逐字校验。第二轮补上之后,数字从 6.19 us 到 6.14 us。早期的 4.7 us 不是假象。它是 A40 上的数,不能拿去和 V100 上的 MSCCL++ 直接比。
从这些测量可以直接得到什么
小消息跨节点通信的瓶颈在端点软件,而且主要在每次操作重复付出的控制步骤上。在无 GDR 的 EDR IB 上,8 B Graph allreduce 13.6 us 里硬件只占 1.15 us,PCIe 往返加 payload 3.6 us,其余约 8.9 us 是代理唤醒、中转与协议同步。CUDA Graph 只解决主机发射,这一块它碰不到。结论在消息低于几 KiB 时成立,到 64 KiB 时差距已缩到 2 倍左右。
没有 GDR 的集群上,新一代设备发起通信栈会整体失效,而且报错不会告诉你是驱动配置问题。NCCL 普通路径会静默退回主机中转,所以日常训练与推理看不出异常;GIN、NVSHMEM、UCCL、DeepEP 则在初始化阶段直接失败。判断这类失败时,先查 nvidia_peermem 是否加载、驱动是否为开源内核模块、CUDA 属性里的 GDR 与 DMA-BUF 支持位,再去讨论机制本身。
一个与你的设计同构、只因环境跑不起来的现成系统,仍然是你的 baseline。MSCCL++ 的失败看上去像环境不支持,实际是一个 buffer 分类细节,两处改动就把它拉到我们的上界附近。如果停在报错那一步,我们会把一个调优问题误读成一个能力缺口。反过来,当现成系统的机制本身就依赖缺失的能力(IBGDA 离开 GDR 就不再是它自己)时,它的缺席只能说明两边都没测,不能说明空间属于新方案。
对小消息 send/recv,NCCL_NET_SHARED_BUFFERS=0 值得先试。它在我们的 V100 对上把 8 B 单程从 23.5 us 降到 11.6 us,比其余所有旋钮的效果加起来都大;allreduce 对它不敏感。节点对之间的差异(同口径 37.7 对 24.3 us)也大于多数旋钮,比较不同方案时必须固定节点对。
常驻化的端到端价值取决于每 step 通信次数与计算量之比,而不是单次延迟省了多少倍。单次省下 2 到 3 倍,落到 70B 张量并行只剩 2% 到 4%,落到小 batch 跨节点专家并行却有 11% 到 23%。评估这类优化时,先按目标部署的并行方式数清每 step 的通信次数与消息大小,再乘以实测的单次差值。
机制已经在,剩下的是调优
按实测曲线做的端到端估算,只有小 batch 跨节点 MoE 专家并行的 decode 能省下 10% 以上。常驻 proxy 这条机制已经在,剩下的是调优。