blog

2026

Hybrid PD Handoff: Does Waiting by Component Beat Waiting by Layer?

On a deterministic model of hybrid-model prefill/decode handoff, waiting per component improves latency by at most 1.98% over waiting per layer, and the median is zero. Streaming by layer, against waiting for the whole transfer, reaches 47.5%.

Agent Sandbox: What Can Still Cut Runtime Memory and Page Faults

On classic LRU, keeping the next call's pages in the compressed pool beats a perfect prefetch. On Linux 6.12, MGLRU plus zsmalloc already cuts stock wakeup reads from 42 MB to 14 MB, and placement no longer wins a stable 20%.

Agent Sandbox Wakeup: What to Prefetch Matters More Than When

A tool call is visible 1.7 seconds before it can be dispatched, but the cost of waking a reclaimed sandbox is which pages come back and how they sit on swap. Reclaiming only the long human-wait gaps already captures nine tenths of the idle memory.

Where Communicator Rebuild Actually Costs

Rebuilding an NCCL communicator on demand takes 50 to 180 ms, almost all of it the first collective's lazy connect. A prewarmed pool of every TP shape costs about 1% of GPU memory and leaves a sub-millisecond switch.

Bit-Exact After the TP Shape Changes?

TP1 has no collective, so bit-exact results across tensor-parallel shapes need one shared tree for the on-GPU K reduction and the cross-GPU reduction. TBIK already builds that tree.

Completion Semantics of NCCL Failure Recovery

On NCCL 2.32.3 send/recv, sender completion is not receiver consumption. Retrying from sender completion loses 1–4 messages or duplicates 4–9; a receiver consumption log drives both to zero.

How Much Memory Can an Agent Sandbox Swarm Still Share?

After a shared read-only repo, only 4% to 10% of a same-repo sandbox swarm is still duplicate, and almost all of that is private heaps, not files.

Can Cross-Node Small-Message Communication Stay Resident?

About two thirds of an 8 B cross-node allreduce is per-operation control cost. A resident host proxy can pay that once, and MSCCL++ already does, within 1.6x of the bound after two patches.

Logical Skew Is Not Physical Contention: How We Killed a Beautiful MoE Systems Idea Early

A negative-result postmortem of IncastEP: skewed MoE routing did not produce multi-source receiver incast under local EP=4 workloads, and the broader lesson is to validate three mappings before building the mechanism.

RL Long Tail: Mid-Batch Parallel Shape Switching

On four PCIe A100s, switching an RL rollout from DP4 to TP4 once the long tail remains is 1.31 times slower than staying on DP2×TP2. The mixed shape wins both phases, and one switch costs at least 19 seconds.

2025

Why Pipeline Parallelism Matters for Serverless LLM Inference

Exploring why pipeline parallelism is critical for LLM inference in serverless environments, where resource constraints and sparse workloads make tensor parallelism impractical, and how pipeline parallelism naturally addresses queuing delays.

FPGA in the AI Era: From Standalone Struggles to Co-Design Opportunities (Part 1)

Exploring FPGA's evolving role in AI infrastructure: why standalone FPGA struggles against GPUs, the limitations of network acceleration, and the promising path forward through FPGA+GPU co-design for pipeline parallelism.

Diffusion Model Serving: Why It Matters for Systems

Exploring why diffusion models matter for AI infrastructure, their fundamental differences from LLMs, and insights from production serving systems. A deep dive into computational patterns, memory subsystems, and the future of diffusion-based LLMs.