Hybrid PD Handoff: Does Waiting by Component Beat Waiting by Layer?

State can be split finely. How much compute can start early depends on the consumer's dependency chain and on the scheduler's boundary. On a deterministic model, waiting per component improves latency by at most 1.98% over waiting per layer, and every workload's median is 0. Waiting per layer, against waiting for the whole transfer, reaches 47.5%.

From moving state to being ready to compute

Large-language-model inference has two shapes of work. Prefill reads the prompt and builds the state later tokens need. Decode reuses that state and emits tokens one by one. Prefill/decode disaggregation, PD disaggregation, puts the two on different GPUs. Each side can then be sized on its own, and the two interfere less.

DistServe picks resources and a parallel strategy for each side from time-to-first-token and time-between-tokens. Handoff bandwidth is part of that placement (Zhong et al. 2024, OSDI). Splitwise starts from the different compute, memory, and power mix of prompt processing and token generation. The two sides may use different hardware, and the intermediate state moves over the interconnect (Patel et al. 2024, ISCA). Separation buys placement freedom. It also forces a question. How much state does decode have to wait for before useful compute can start?

On an attention layer the object is the KV cache, the keys and values of past tokens. It grows with the prompt. A hybrid model mixes attention with a linear recurrent layer, so the state is more than KV. A Gated DeltaNet or Mamba-style layer keeps a fixed-size recurrent state and the tail of a short convolution window. Below, ssm is the recurrent state and conv is the convolution tail. The names classify handoff pieces. They leave open whether the models share one recurrence.

Prefill layers produce KV, recurrent state, and convolution tails. Decode consumes them along a per-layer chain.
Figure 1. Prefill on top, the state that has to move in the middle, decode below. Blue KV grows with the prompt. Orange recurrent state and green convolution tails stay fixed size. Red arrows are the decode layer chain. More component types do not add a branch around that chain.

The same recurrent layer’s ssm and conv both serve one decode layer. The next layer still waits on that layer’s output. One narrow case remains. If conv arrives first and the runtime can split the work inside a layer, the convolution operator can start. The recurrent layer is still unfinished. A later layer still waits. Component-level readiness buys only a short head start inside the layer.

Arrived, runnable, and admitted are different events

A PD handoff has a data path and a compute path. The producer has to know when a layer’s state can be sent. The consumer has to know when the bytes on hand are enough for the next piece of compute. The scheduler still decides when the request may enter a batch.

vLLM’s disaggregated-prefill connector stores and loads state with the attention module, layer by layer. save_kv_layer is the producer’s per-layer send point. wait_for_layer_load is the consumer’s per-layer wait (vLLM disaggregated prefill). LMCache stores and moves KV. A layer-wise integration can overlap the load with compute (LMCache). Per-layer streaming already has an interface. NIXL moves bytes point to point across memory and storage backends (NIXL). A transport library moves bytes. When a transfer counts as done, which pieces are waited on, and which request may run, stay with the layer above. The library in use leaves the readiness granularity open.

Schematic timelines for waiting on the whole transfer, waiting per layer, and waiting per component. The first-token line barely moves between the last two.
Figure 2. A schematic layer-time axis. Gray is prefill, color is the transfer, black is decode, and the red line is first-token completion. Global wait still blocks the consumer until every layer has arrived. Per-layer readiness lets early decode layers run during later transfers. Per-component readiness barely moves the red line.

Removing the whole-request barrier releases several layers of compute. Splitting components inside a layer pays only when those components arrive apart, and when an operator can run on the piece that arrived. Figure 2 sketches that mechanism. Later figures carry the model latencies. Overlap differs by workload.

TL;DR

The question is whether a hybrid-model handoff needs a dependency graph finer than the layer. On a single-request model, component readiness beats layer readiness by at most 0.053%. The heavier sweep has 648 deterministic replay points. It covers speculative decoding, continuous batching, interleaved requests, and pipeline parallelism. The maximum there is still 1.98%. Every concrete workload’s median is 0.00%. Layer readiness against a global wait reaches 47.5%.

All of these are model values. The gain depends on whether the consumer has independent work it can start, and on whether the scheduler allows that work to run. More components do not by themselves mean more usable parallelism.

The hypothesis

The component story is specific. Put KV, ssm, and conv on a directed acyclic graph. Run a node as soon as its dependencies are met. That should release more compute than ordinary per-layer readiness. The ready frontier is the set of nodes whose dependencies are already satisfied.

The increment that matters is component-level against layer-level. Per-layer streaming can take the same gain over a global wait, so a win against global wait leaves the component question open. A 15% latency drop is the line used here for a gain worth caring about. The comparison uses the whole distribution.

A single-request layer chain is a weak test. The stronger test is a heavier workload. Speculative decoding adds a draft and a target. Continuous batching runs many requests together. Interleaving shares one link across requests. Pipeline parallelism cuts the model across devices. The question is whether those change what the consumer depends on, or only how the same chain queues and where it is placed.

Same bytes, different waits

The model covers Qwen3.6-35B-A3B and Nemotron-3-Nano-30B-A3B. Prompts are 1k, 8k, and 32k. Links are PCIe Gen4, 100G, and 25G. PCIe Gen4 bandwidth is a measured input. The handoff delays and the policy gains computed from it are still model outputs.

Prefill model FLOPs utilization, MFU, is 0.4 or 1.0. The first is a lower efficiency. The second is idealized. Two timings stay apart. Handoff may send while prefill is still computing. Migration starts the move only after compute finishes. Their overlap differs, so they stay separate configurations.

The four policies share the same bytes and the same compute. A finishes compute, then moves state. B is the stronger global-wait baseline. When streaming is allowed, the producer already sends layer by layer, and the consumer waits for every piece. C is ready per request and per layer. A decode layer waits for that layer’s state and for its compute predecessor. D splits the components inside C, and may run the part of the layer whose inputs have arrived.

The single-request sweep is 108 configurations and 216 comparisons. The event-graph sweep over four workload families is 648 replay points. A gain is the latency drop against the named baseline. Boxes and points are differences across configurations. Read them as configuration spread. Run-to-run noise and confidence intervals are outside this sweep.

The model takes a prefill MFU, an attention shape, a convolution operator share of 10%, and 0.05 ms to move an activation between pipeline stages. Request-level asynchronous execution assumes zero interference, and it leaves out the extra weight reads from splitting a batch. Semantic correctness of the schedule sits outside this model.

Short prompts still move a fixed state

Handoff bytes per request at three prompt lengths. Recurrent and convolution state stay flat. KV grows with the prompt.
Figure 3. Bytes moved per request. Orange is recurrent state plus the convolution tail. Blue is KV. At 1k, Qwen3.6 moves 85.4 MB, about 64 MB of it recurrent, 75%. Nemotron moves 55.4 MB, 89% recurrent. At 32k the totals are 736 MB and 250 MB, and the recurrent share falls to 9% and 20%.

A shorter prompt leaves the recurrent state the same size. That fixed state is a floor under the transfer, so a short request can still pay a real move. At long prompts the recurrent share falls because KV grows. Attention keeps a KV entry per past token. A recurrent layer compresses history into a fixed shape. The orange bars stay nearly flat. The blue bars grow. That mix is why a short prompt on a slow link is where streaming matters. Bytes set how long the move takes. The consumer’s dependencies set what can run during the move. A pure attention model has no recurrent floor. Partial reuse would change which bytes move. The share of recurrent bytes and the share of end-to-end latency are different quantities.

The single-request gain is the layer stream

Single-request time-to-first-token gains for Qwen3.6 at 1k. Per-layer streaming against compute-then-move is large. Per-component readiness against per-layer readiness is nearly invisible.
Figure 4. Qwen3.6, 1k prompt, one request. Blue is per-layer readiness against compute-then-move. Orange is per-layer against a global wait whose producer already streams. Green is per-component against per-layer. The green bars are almost absent.

On 25G, per-layer readiness beats compute-then-move by up to 47%, and still by 31% at MFU 0.4. Against a global wait whose producer already streams, only the MFU 1.0 and 25G point clears 15%, at 16.5%. At MFU 0.4 that increment is 0. Component against layer improves by at most 0.053%. That is why the green bars are almost absent. Moving from compute-then-move to per-layer streaming opens overlap on both sides. Producer compute overlaps the transfer, and consumer compute overlaps it too. The global-wait baseline has already taken the producer half. What remains is idle time on the consumer. The tall blue bars are the layer stream.

conv and ssm of one layer come from the same prefill layer. Even if the convolution tail is sent first, decode can run only the convolution. It still needs the recurrent state to finish the layer. Qwen3.6’s convolution tail is about 49 KB per layer. On 25G that lead is about 16 µs. The overlap window is that arrival gap. The volume of the recurrent state is a different quantity. This is a single-request layer chain. 0.053% is the maximum increment in these configurations. When producer and consumer are both organized by layer, a component signal adds little beyond per-layer streaming.

Four workloads do not widen the per-request frontier

Speculative decoding drafts candidate tokens and has the target model check them, so fewer expensive serial steps are needed (vLLM speculative decoding). The model here is one draft layer and a candidate length k=3. That is the setting, and other speculative algorithms sit outside it. Continuous batching puts active requests in one decode step and updates the set on the step boundary. Interleaving is about how requests share the state link. Pipeline parallelism, PP, places layers on different devices and forwards activations.

Event dependencies for speculative decoding, continuous batching, interleaved requests, and pipeline parallelism. Extra chains and couplings do not add a branch around the layer predecessor.
Figure 5. Boxes are compute events. Arrows are order. Colors separate requests or stages. This is not a timeline to scale. The workloads add chains, coupling, or a device boundary. They do not add a component branch that skips the previous layer.

In the speculative panel, the modeled draft depends on the target’s last hidden state, so it hangs off the end of the target chain. The extra model role stays off the target’s handoff path. In the continuous-batch panel, dashed lines are requests sharing one GPU operator for a layer. The requests keep separate state. The runtime still binds them, which is why a whole batch can stall together. In the interleaved panel each request keeps its own layer chain. A shared link changes arrival order. The PP panel cuts one chain across devices and adds an activation hop. The later stage still needs the earlier stage’s output. Contention and cross-device traffic move the timeline. Components of one layer stay consumed together. Tree speculation, EAGLE-3’s multi-layer features, chunked prefill, and a partial handoff after a prefix-cache hit are outside this figure.

Layer streaming takes the large gains

Maximum time-to-first-token gain by workload and link. Per-layer readiness against a global wait is large. Per-component readiness against per-layer readiness stays near zero.
Figure 6. Each bar is the maximum inside its group, not a typical request. Orange is per-layer against a global wait. Green is per-component against per-layer.

Per-layer readiness improves speculative decoding by up to 47.5%, continuous batching by about 37% to 38%, interleaving by about 23% to 24%, and PP by about 45%. The component increment tops out near 2.0%, on continuous batching at 25G. PP is about 0.8% to 0.9%. The rest sits near 0. Heavier workloads widen some gaps. The wide gap is still whether the consumer may run layer by layer. Splitting that chance into components inside a layer adds almost nothing. A global wait holds runnable compute behind one all-done event. Per-layer readiness releases it. The component frontier remains the per-request layer chain. MFU 0.4 and 1.0 look alike, so the small component increment is common to both efficiencies. Each bar is a group maximum, and the maxima need not share one configuration. The 47.5% is layer streaming on speculative decoding. It stays separate from the component graph.

Across the distribution, the component frontier stays flat

At MFU 0.4, the distribution of gains. Component against layer sits on zero. Layer against global wait and request-level async execution do not.
Figure 7. MFU 0.4. Red is component-level against the strongest per-layer policy, and it sits on zero. Gray is per-layer against a global wait. Blue is request-level asynchronous execution against the per-layer policy.

On continuous batching the gray median is about 12% and the maximum about 37%. Blue is request-level async execution. On that family its median is about 23% and its maximum is 88%. State can be ready while the request is still waiting to run. The blue box is execution freedom. An independent request need not wait out the original batch. That freedom and a finer state graph are different questions. Component readiness changes how a layer’s pieces are waited on. Request-level async changes how requests are bound together. The second one touches the step boundary and the shared operator, so it can move a longer wait. 88% is an upper bound under zero interference. The bound leaves out the extra weight traffic of a split batch. Splitting a batch changes utilization, memory traffic, and launch cost. The boxes are sensitivity across configurations.

Component-level gain over per-layer readiness across 648 replay points. Medians are zero. The largest point is about 2%.
Figure 8. The vertical axis is the component increment alone. c is requests in the batch. K is interleaved requests. Every group's median is 0.00%.

Across all 648 points the median of every group is 0.00%. The speculative maximum is 0.15%. Batch sizes 8 and 32 reach 1.97% and 1.98%. Interleaving at 4 and 16 reaches 0.26% and 0.20%. PP2 is 0.27% on a private link and 0.85% on a shared link. PP4 is 0.23% and 0.70%. A small change in arrival time can push a request across an execution boundary, so a few points leave zero. The 1.98% maximum sits on that kind of threshold. Reading it as general component parallelism of that size would stretch one boundary effect. Give the per-layer policy the same request-level async freedom, and the gap between the two is at most 0.34%. A fair comparison has to match both the readiness rule and the right to run. The result is for the models and schedulers in this sweep.

The step boundary is the wait that moves

A new request in continuous batching usually cannot join a step that is already running. Even if its state has arrived, it waits for the next admission point and then does its own compute. Discrete steps turn a continuous arrival into a whole-step wait. If the arrival phase is roughly uniform and a batched decode step takes D_c, the expected time from state-complete to the end of an execution step is about D_c/2 plus D_c. The first term is the wait for the boundary. The second is the step itself. In the model D_c is about 23 to 38 ms. A 1k transfer is about 7 to 27 ms. The scheduler boundary can already outweigh the transfer.

A request that arrives mid-step still waits for the next step. Per-component readiness matches per-layer latency. Admitting too early stalls the whole batch.
Figure 9. Left, state that arrives mid-step still waits. Center, first-token latency for a new request. Right, extra delay imposed on requests already running if an unready request is admitted into the batch. The right axis is logarithmic.

The center panel is Qwen3.6, 1k, 25G, MFU 0.4, handoff. For batch sizes 8 and 32, global wait is 88 ms and 112 ms. Per-layer readiness is 65 ms and 72 ms. Per-component readiness is still 65 ms and 72 ms. Request-level async execution is 53 ms and 53 ms. Finer components save nothing in this example. Async execution recovers 12 to 19 ms, and that recovery is the step-boundary quantization. While the request sits on an admission instant, a more precise state signal leaves the GPU on the same schedule. Shortening that wait means changing when a request may join, and how it runs with the others.

The right panel is the other failure. Admitting too early can hold the requests already running. In the vLLM per-layer hook modeled here, wait_for_layer_load waits on the whole batch and a layer name. Other requests in the batch have no bypass. Put a request that is still handing off into the batch. When execution reaches that layer, the whole batch can block on it. At 1k the extra delay is 17.5 ms at batch 8 and 2.0 ms at batch 32. At 8k it is about 0.39 to 0.42 seconds. At 32k it is about 2.28 to 2.30 seconds. At a long prompt the delay approaches the prefill itself. The consumer is waiting for state the producer has not created yet, on top of bytes already in flight. A per-layer signal should let useful compute proceed, and one unready request should stay off the batch’s wait condition. The hook modeled here hangs only on attention layers. GDN and Mamba layers have no matching hook. A per-layer KV interface covers KV. The other hybrid handoff states are separate. The right panel follows from the hook and from the admission assumption. It is a model of that hook. Request-level admission changes the behavior. A real async gain also pays for splitting the batch and for contention.

The consumer chain sets the gain

The opening question comes from the state mix. A hybrid model hands off KV, a recurrent state, and a convolution tail, so a component graph looks natural. In the models and four workloads covered here, each request still advances along a layer chain. Components of one layer are produced together and consumed together. Splitting them leaves a narrow frontier.

The gains that show up are per-layer streaming against a global wait, and the scheduling room that appears once requests may run apart. Those answer layer-to-layer waiting and request-to-request binding. Component granularity is a separate increment.

Every number is a deterministic model output. Semantic correctness sits outside this model. Prefill efficiency, attention shape, the draft, and pipeline transfer are assumptions. Raising the convolution operator’s share from 10% to 50% still leaves the component increment bounded by how far ahead the convolution state arrives. Multi-layer feature dependencies, tree speculation, and a partial handoff sit outside the model. The conclusion stays inside the cases above.

What a finer ready signal can and cannot buy

The value of a finer ready signal is set by what the consumer depends on. If components are produced together and merged by one consumer, a split completion signal can use only the real gap between their arrivals. Extra kinds of state become extra runnable work only when that gap is real.

A handoff optimization has to keep the overlap the system already has inside the baseline. Compute-then-move, a global wait with a streaming producer, and per-layer readiness are different increments. If the producer can send early, that ability stays in the baseline when a finer mechanism is judged. Leave it out, and the gain lands on the wrong change.

More requests and more devices add coupling. One request’s dependency frontier is a separate question. A shared operator, a shared link, and a pipeline cut all move where the wait sits. Whether state can be consumed early is still a data dependence, checked edge by edge.

After the data is ready, the scheduler’s quantization can be the larger wait. In the continuous-batch model the step and the transfer are the same order of time, and the ideal gain of request-level async execution is much larger than the component split. Admission that happens only on a discrete boundary has to be counted apart from state wait and from compute. The async upper bound also has to be kept apart from the cost of splitting the batch.

The scope of a wait hook is part of the performance semantics. If a per-layer wait blocks the whole batch, one unready request passes the producer’s delay on to requests that were already running. The bytes being waited on are one question. Who else waits, and whether every state type has a hook, are the next ones.

RL Long Tail: Mid-Batch Parallel Shape Switching asks whether an RL rollout should change parallel shape in the middle of the batch. It does not measure a prefill/decode split. This post starts after that split already exists, and asks how fine the handoff’s ready signal has to be.