blog
2026
On a deterministic model of hybrid-model prefill/decode handoff, waiting per component improves latency by at most 1.98% over waiting per layer, and the median is zero. Streaming by layer, against waiting for the whole transfer, reaches 47.5%.
On classic LRU, keeping the next call's pages in the compressed pool beats a perfect prefetch. On Linux 6.12, MGLRU plus zsmalloc already cuts stock wakeup reads from 42 MB to 14 MB, and placement no longer wins a stable 20%.
A tool call is visible 1.7 seconds before it can be dispatched, but the cost of waking a reclaimed sandbox is which pages come back and how they sit on swap. Reclaiming only the long human-wait gaps already captures nine tenths of the idle memory.
Rebuilding an NCCL communicator on demand takes 50 to 180 ms, almost all of it the first collective's lazy connect. A prewarmed pool of every TP shape costs about 1% of GPU memory and leaves a sub-millisecond switch.
TP1 has no collective, so bit-exact results across tensor-parallel shapes need one shared tree for the on-GPU K reduction and the cross-GPU reduction. TBIK already builds that tree.
On NCCL 2.32.3 send/recv, sender completion is not receiver consumption. Retrying from sender completion loses 1–4 messages or duplicates 4–9; a receiver consumption log drives both to zero.
After a shared read-only repo, only 4% to 10% of a same-repo sandbox swarm is still duplicate, and almost all of that is private heaps, not files.
About two thirds of an 8 B cross-node allreduce is per-operation control cost. A resident host proxy can pay that once, and MSCCL++ already does, within 1.6x of the bound after two patches.
A negative-result postmortem of IncastEP: skewed MoE routing did not produce multi-source receiver incast under local EP=4 workloads, and the broader lesson is to validate three mappings before building the mechanism.
On four PCIe A100s, switching an RL rollout from DP4 to TP4 once the long tail remains is 1.31 times slower than staying on DP2×TP2. The mixed shape wins both phases, and one switch costs at least 19 seconds.
2025
Exploring why pipeline parallelism is critical for LLM inference in serverless environments, where resource constraints and sparse workloads make tensor parallelism impractical, and how pipeline parallelism naturally addresses queuing delays.
Exploring FPGA's evolving role in AI infrastructure: why standalone FPGA struggles against GPUs, the limitations of network acceleration, and the promising path forward through FPGA+GPU co-design for pipeline parallelism.
Exploring why diffusion models matter for AI infrastructure, their fundamental differences from LLMs, and insights from production serving systems. A deep dive into computational patterns, memory subsystems, and the future of diffusion-based LLMs.