Logical Skew Is Not Physical Contention: How We Killed a Beautiful MoE Systems Idea Early

0. MoE communication in two figures

Before discussing the failed idea, it helps to be precise about what MoE actually communicates. In one MoE layer, the router does not compute the expert output; it selects a small set of experts for each token. Those experts run their MLPs, and their outputs are combined back at the token owner. MoE therefore creates a sparse token→expert dependency graph, not a dense all-to-all semantic dependency.

A MoE layer: tokens are sparsely routed to experts, then expert outputs are combined
A MoE layer: tokens are sparsely routed to experts, then expert outputs are combined

With Expert Parallelism (EP), experts are placed across GPUs. A token owned by one GPU may need to cross NVLink, PCIe, or RDMA to reach an expert on another GPU; the expert result then returns for combine. Every MoE layer therefore induces a dynamic source GPU → destination rank traffic matrix.

Expert Parallelism maps logical token→expert routing onto physical GPU→GPU dispatch and combine
Expert Parallelism maps logical token→expert routing onto physical GPU→GPU dispatch and combine

The subtle point—and the starting point of this story—is that the model’s “hot expert” graph and the network’s “how many independent sources converge on one physical receiver” graph are not the same graph. An expert may be hot while most of its tokens come from one source. A destination rank may host many distinct expert flows while still seeing only one or two senders.

Our original IncastEP idea made a plausible but unverified jump exactly here: routing skew → hot expert → many-source incast. The rest of the blog is the story of how real workloads dismantled that chain one mapping at a time.

TL;DR. We started from a compelling hypothesis: skewed MoE routing should create many sources → hot expert traffic, which should produce receiver-side incast; an expert-aware communication scheduler could then prioritize transfers that unlock useful computation and reduce MoE layer latency. The chain was elegant—but the workload did not cooperate. Across local EP=4 runs of GLM, Qwen, and Nemotron, both expert-level and destination-rank source fan-in stayed at p95=1 and max=2. Nemotron could stack many distinct expert flows on one destination, but they still came from only one or two source GPUs. Meanwhile, calibrated expert MLPs took only about 0.07–0.20 ms, leaving a much smaller “unlock compute early” window than our original synthetic setup suggested.

We stopped before building credits, adaptive admission, or a custom kernel. The broader lesson is more useful than the abandoned design: before optimizing a bottleneck, validate three mappings—application semantics → physical resource contention, contention → critical-path delay, and experimental policy → actual hardware behavior.

The hypothesis chain and where it broke
The hypothesis chain and where it broke

1. Why the idea looked so plausible

MoE makes a seductive systems story. The router assigns tokens to experts, expert popularity is not perfectly uniform, and some experts become hotter than others. It is natural to imagine that a hot expert attracts traffic from many GPUs at once. At the network layer, that looks like classic many-to-one incast; at the model layer, the same hot expert can look like a compute straggler.

That gives a clean causal chain:

Routing skew
    ↓
Hot experts
    ↓
Many sources → one receiver
    ↓
Receiver-side incast
    ↓
Expert-aware communication scheduling
    ↓
Prioritize flows that unlock useful GEMM
    ↓
Lower MoE layer latency

It also suggests an attractive research tension: network-optimal need not be model-optimal. Suppose Expert A needs only four more tokens to launch its next useful GEMM chunk, while Expert B already has a full chunk waiting in its compute queue. A network-centric scheduler might prefer a larger transfer to B because it packs the link well or fits a matching objective. A model-progress scheduler might send A’s four tokens first, start useful compute earlier, and overlap it with the rest of the communication.

The reasoning is not logically wrong. The problem is that it silently crosses three unverified mappings:

  1. Does a hot expert actually imply many independent physical sources converging on one receiver?
  2. If arrival order changes, is the unlocked expert compute long enough to move the layer critical path?
  3. Does the experiment’s “different schedule” actually change what is launched on the wire?

Our experiments ended up falsifying—or at least sharply weakening—each assumption in turn.


2. First break: logical skew did not become multi-source physical incast

We eventually stopped looking only at “hot experts” and measured two distinct fan-in metrics.

For an expert e:

\[F_e = \left|\{s\mid \exists\ t\text{ from source }s\text{ routed to }e\}\right|.\]

For a physical destination rank r:

\[F_r = \left|\{s\mid \exists\ e\mapsto r,\ s\text{ sends remote tokens to }r\}\right|.\]

The first is expert-level fan-in. The second is destination-rank fan-in. This distinction matters because low per-expert fan-in does not rule out physical contention: many different experts might be placed on the same rank or behind the same NIC, causing their traffic to converge again at the receiver.

We increased concurrency on GLM-4.7-Flash from 4 to 32, checked Qwen3.6-35B-A3B, and then used Nemotron-3-Nano at concurrency 32/64/128 as a stronger-expert and richer-flow case.

The result was stubbornly flat: expert fan-in p95=1, max=2; destination-rank fan-in p95=1, max=2. Increasing request concurrency did not push source cardinality toward the 8–16-source regime that the original IncastEP story implicitly relied on.

Expert and destination-rank fan-in remain flat across the measured configurations
Expert and destination-rank fan-in remain flat across the measured configurations

This became the most important negative result in the project. Under these local EP=4 configurations, the arrow

routing skew → many-source receiver incast

was simply not supported.

The subtle trap: many flows are not many sources

Nemotron made the distinction especially clear. At concurrency 128, a destination rank could see roughly p95=4, p99=12, and max=23 distinct expert flows. A message-count proxy reached about 30 in the tail. Looking only at the number of logical flows, the receiver appears busy and fragmented.

But plot those flow counts against the number of unique remote source GPUs for each (dispatch, destination rank):

Nemotron c128: many logical expert flows can still come from only one or two source GPUs
Nemotron c128: many logical expert flows can still come from only one or two source GPUs

Almost the entire distribution sits on one or two sources.

A receiver can carry many semantically distinct expert flows without experiencing multi-source incast. Logical fragmentation is not physical fan-in.

That distinction is deeper than this one project. Models expose tokens, experts, top-k dependencies, and activation histograms. NICs expose endpoints, QPs, queues, bytes, and concurrent senders. Those are different graphs.

We do not have enough evidence to claim a universal reason for the persistent 1–2-source pattern. Fine-grained expert structure, token ownership, placement, and the serving runtime’s dispatch organization may all contribute. The important methodological point is not to invent a mechanism explanation after the fact. The measurement we can defend is narrower: the multi-source convergence assumed by our design did not appear in the configurations we measured.


3. Second lesson: a microbenchmark can make a bottleneck real without making it prevalent

Earlier synthetic experiments looked encouraging. In a cross-node, fixed-total-byte N→1 test at roughly 344 KiB offered traffic, increasing fan-in from 1 to 3 increased median network completion from roughly 0.29 ms to 0.81 ms.

Synthetic N-to-1 traffic slows with fan-in, while the real traces stayed in the 1–2-source region
Synthetic N-to-1 traffic slows with fan-in, while the real traces stayed in the 1–2-source region

That observation was not wrong. If we force several sources to hit one receiver, multi-source contention can become a real network effect. But a microbenchmark answers a mechanism question, not a prevalence question.

The two responsibilities should be separated explicitly:

  • Microbenchmarks establish causality: if fan-in increases, do arrival spread, receiver pressure, or completion time worsen?
  • Workload characterization establishes prevalence: how often do real model executions actually enter that fan-in/flow-size regime?

Running only the first can produce a dangerous kind of paper: a system that successfully fixes a bottleneck that its benchmark manufactured.

One of the most useful process changes in this project was therefore to require controlled workloads to be calibrated from measured traces—source cardinality, flow sizes, expert mix, and compute cost—rather than designing a beautiful W1/W2 first and then searching for a workload that resembles it.


4. Third constraint: even a correct dependency may live on the wrong timescale

The original idea also relied on a second assumption: once an expert becomes ready earlier, its compute is long enough to overlap meaningful remaining communication and move the layer completion time.

Our first synthetic setup used large GEMMs, which made this easy to demonstrate. Calibrating the expert MLP with real model dimensions changed the picture.

On an A40:

  • GLM-4.7-Flash was around 0.07 ms for 1–32 tokens and about 0.095 ms at 256 tokens;
  • Nemotron-3-Nano was around 0.08 ms at small token counts and about 0.12 ms at 256;
  • a DeepSeek-V4-Flash proxy calibration was around 0.11–0.12 ms for small batches and about 0.20 ms at 256 tokens.
Calibrated expert MLP latency is only about 0.07–0.20 ms in these experiments
Calibrated expert MLP latency is only about 0.07–0.20 ms in these experiments

A useful first-order bound is:

\[\Delta T_{layer} \lesssim \min(T_{GEMM\ unlocked}, T_{comm\ remaining}) - T_{scheduling\ overhead}.\]

If the expert compute itself is only 70–200 microseconds, the maximum useful overlap window is already small. Add admission logic, queue management, synchronization, or kernel-launch overhead, and a semantically correct scheduling improvement can quickly collapse into the noise floor.

The lesson is not that communication-compute overlap is useless. It is more precise:

Criticality in a dependency graph is not enough; the dependency must also have enough wall-clock budget to justify a separate mechanism.

This is why we stopped using inflated GEMMs as scientific evidence. If a scheduling idea requires multiplying the real expert compute time by ten to pass a Go threshold, we are validating the synthetic regime, not the original systems claim.

The DeepSeek-V4 serving trace itself was blocked by the current A40/vLLM FP8 KV-cache requirement, so the DeepSeek curve here is calibration-only, not end-to-end serving evidence.


5. A different kind of failure: our first negative experiment did not actually change the network schedule

One of the most useful lessons came from the experimental harness rather than the workload.

In the first W1 implementation, the receiver posted both irecvs up front and both senders immediately issued isend. What we called network_oracle and model_oracle mostly changed the receiver’s wait() order and when GEMM was launched.

We thought we were comparing:

Schedule A: flow 1 launches first, flow 2 later
Schedule B: flow 2 launches first, flow 1 later

but both flows were already live on the wire. We were comparing software observation order, not communication admission order.

Phase 2a fixed this with grant-before-isend. The receiver maintains an admission budget; a sender cannot issue the payload until it receives a grant. Only then can S_net=[2,1] and S_model=[1,2] produce genuinely different launch timelines.

The corrected harness changes send-launch order on the wire
The corrected harness changes send-launch order on the wire

In a deliberately asymmetric harness-validation case, S_net=[2,1] achieved T_net≈0.410 ms, while S_model=[1,2] produced T_layer≈1.252 ms; the oracles disagreed, and the measured schedule gap was about 0.027 ms. But this case intentionally used highly asymmetric GEMMs (M=4096 vs M=8). It proved only one thing:

The harness can now instantiate two different physical schedules and can demonstrate that network and model objectives may disagree in a constructed case.

It did not prove that a useful disagreement exists in a real MoE workload.

This gives a broadly applicable experimental rule:

Before interpreting a negative result, verify that the two policies actually produce different physical execution at the system boundary.

For a network system, inspect grant/launch/packet timelines. For a GPU runtime, verify kernel order. For storage, verify the actual I/O path. A variable named network_oracle is not evidence that the network did something different.


6. The real lesson: validate three mappings before designing the mechanism

The IncastEP exploration can be compressed into three checks that are useful for almost any systems idea.

Mapping 1: application semantics → physical contention

Skew, hotness, sparsity, or dependencies at the model level do not automatically imply incast, queue buildup, or link imbalance at the network level.

Our concrete counterexample was:

many/hot expert flows
      ≠
many unique source GPUs
      ≠
receiver-side incast

Measure the endpoint-level traffic graph before designing the transport.

Mapping 2: physical contention → critical-path budget

Even if contention is real, ask whether it sits on the metric that matters and how many microseconds are actually recoverable.

Here, expert compute was often around 0.1 ms, so “make this expert ready earlier” had a very limited budget to turn into end-to-end layer latency.

Calibrate real compute and communication times before reasoning about overlap.

Mapping 3: experimental policy → actual hardware behavior

Changing a queue in software does not guarantee that the NIC saw a different schedule. Changing a wait order does not mean send order changed.

Every hypothesis test needs a sanity check that the policy actually changed the machine.

We now like to express this as a short pre-mechanism checklist:

1. Am I measuring a model-level metric or a physical-resource metric?
2. How much wall-clock critical-path budget does this physical effect occupy?
3. Do the compared policies actually differ on the wire/GPU/I/O path?
4. Does the synthetic stress regime fall inside the real workload distribution?
5. If the real workload does not enter that regime, am I willing to stop?

The fifth question is the hardest—and probably the most valuable.


7. Does this prove that MoE never has incast? No.

The result must be scoped carefully.

What the data supports is:

In the local EP=4 continuous-batch configurations we measured for GLM, Qwen, and Nemotron, we did not observe the strong multi-source expert/destination-rank incast required by the original IncastEP story.

It does not establish that:

  • all MoE models avoid incast;
  • EP=8/16, larger clusters, different token ownership, or different dispatchers behave the same way;
  • coarse-grained 8/16-expert models cannot produce stronger fan-in;
  • production-scale NIC congestion is absent;
  • many small expert flows have no optimization opportunity.

We deliberately did not download Mixtral just to rescue the story. Cross-node EP=8 was skipped after the local destination-rank signal stayed flat. DeepSeek-V4 serving traces were blocked by current FP8 KV-cache support. And the message counts in our routing analysis are (src,dst,expert)-level proxies, not NIC hardware counters.

A useful negative result has value because its boundary is explicit, not because it is promoted into a universal theorem.


8. What became more interesting after the idea died?

Stopping IncastEP did not make the data useless. It exposed two questions that fit the measured workload better.

8.1 Token-ready, not expert-ready

With top-k routing, a token cannot enter combine until its required expert contributions are available. The most valuable transfer may therefore be the one that converts many tokens from “partial” to combine-ready, not the one that makes an expert receive a few more tokens.

A new scheduling question would be:

Which pending communication completes the most token dependencies on the critical path?

That is a token-level dependency problem, not an incast problem. It needs a fresh falsification step; it cannot inherit IncastEP’s motivation by assumption.

8.2 Many flows from few sources: fragmentation and coalescing

Nemotron showed p99≈12 and max≈23 distinct expert flows on a destination rank while source fan-in remained 1–2. That points toward a different communication cost model:

  • launch/notification/completion overhead for many small expert messages;
  • coalescing several expert flows on the same src→dst path and scattering at the receiver;
  • metadata and polling overhead that may dominate tiny payloads.

This is almost the opposite of the original design. Instead of throttling many senders, the right primitive may be to coalesce many semantic flows from the same sender.

The useful shift is that the workload is now telling us what mechanism to consider, instead of the mechanism telling us what workload to look for.


9. Stopping early is progress

The dangerous failure mode in systems research is not a null result. It is building months of mechanism before discovering that the motivation never existed in the workload.

We were fortunate to put workload characterization, GEMM calibration, and harness validation in front of credits, adaptive scheduling, and kernel work. The outcome was that we stopped.

That stop saved engineering time and left behind a more reusable principle than the original design:

Before optimizing a system bottleneck, verify application semantics → physical resource contention, contention → critical-path delay, and experimental policy → actual hardware behavior.

Or more plainly:

Do not infer network congestion from model skew, do not infer optimization value from DAG criticality, and do not infer hardware behavior from the names of two software policies.

A beautiful systems story is worth pursuing. A disciplined process that can kill it early is worth even more.