Agent Sandbox: What Can Still Cut Runtime Memory and Page Faults
On a classic LRU, keeping the next call's pages in the compressed pool beats knowing where they are. On Linux 6.12, with MGLRU and zsmalloc, stock reclaim already does most of that. Wakeup disk reads fall from 42 MB to 14 MB, and what remains is not enough for a placement change to win another 20%.
A sandbox for an agent is starting to look like a long-lived process. A language server, a Python tool set, or a headless browser sits idle for seconds to minutes between calls, waiting on the model or on a person. The host wants more of them on one machine, so it reclaims their memory while they wait, and the next call pays to wake them up. This article is about where that wakeup spends its time, and why a placement idea that looks natural has almost no room left on a newer kernel. After reading it you should be able to tell whether a placement or prefetch idea for post-reclaim restore is worth building on the kernel you actually run.
After reclaim, a call has to read pages back
When Linux reclaims an anonymous page, it writes the page to a swap device and leaves a swap slot number in the page table. The next access faults, and the kernel reads the page back by that slot. A 4 KiB read on a SATA SSD is tens to more than a hundred microseconds, so the wakeup cost is roughly the number of pages read back, times the cost of one read, plus the fault handling itself. After a full reclaim, one call into a tsserver sandbox usually reads 50 to 70 MB back, a bit over ten thousand pages.
zswap adds a layer on that path. It is a compressed swap cache in the kernel. A page is compressed into an in-memory pool before it is written to the swap device, and only when the pool is full does the oldest entry get written out, in LRU order (Linux kernel documentation, zswap). The pool size is capped by max_pool_percent, often a few percent of memory. A page that is still in the pool costs a decompression of a few microseconds. A page that has left costs a disk read. When the pool cannot hold the whole sandbox, which pages stay in it decides how much disk the wakeup reads. Kernels since 6.8 also give each cgroup memory.zswap.max and memory.zswap.writeback, so a cgroup can be capped in the pool and can be forbidden from writeback, and a shrinker can push cold entries to disk under memory pressure.
Entries in the pool live in an allocator, the zpool. zbud stores at most two compressed objects in one physical page. It is simple, and the density is low. zsmalloc packs objects by size class, and the density is much higher. The same 5% pool holds many more pages with zsmalloc. A lot of 6.1 configurations defaulted to zbud. The 6.12 default is zsmalloc.
The other variable is the order in which reclaim picks pages. Classic LRU has two lists, active and inactive. A memory.reclaim that drains a whole cgroup toward zero does not pick pages according to whether the next call will use them (Linux cgroup v2 documentation). Multi-gen LRU, MGLRU, sorts pages into generations by age and reclaims from the oldest generation. A page that is accessed is promoted to the youngest generation when generations age. The two youngest generations are the equivalent of the active list, and they are not evicted first (Linux kernel documentation, Multi-Gen LRU, design notes). MGLRU landed in 6.1. Many distribution configurations turn it on by default in 6.12.
Prefetch and a full snapshot already exist
One existing move prefetches on the restore side. When the call arrives, and before the faults, it issues reads for the pages it expects to need, for example with MADV_WILLNEED. It does not move the pages. It turns serial faults into parallel reads. Its ceiling is a prefetch set that matches the set this call will touch. What remains is the cost of reading those pages from wherever they already sit.
The other move is a whole-machine snapshot. Firecracker is a lightweight hypervisor built for serverless, used by AWS Lambda and Fargate (Agache et al. 2020, NSDI). A snapshot is a guest-memory file plus a microVM state file. It can be a full copy or a diff of dirty pages. On restore, the kernel can fault the memory file in, or a userfaultfd handler in userspace can serve it (Firecracker snapshot support). When pages come in on fault, REAP measured snapshot start as 95% slower than a resident start, on average. The same function touches a stable working set across calls, so REAP records that set on the first call, writes it as one compact file, and reads it back in one shot on later restores. Cold start drops by 3.7 times (Ustiugov et al. 2021, ASPLOS). A snapshot costs the host no memory while the sandbox is idle. Every pause has to write the memory state out.
Both moves rely on a stable working set across calls. On these sandboxes that holds. Adjacent calls overlap by 0.82 to 0.84 for tsserver, 0.94 for the Python tool set, and 0.996 for a torch inference sandbox. A JVM workload with random access overlaps by 0.11, and it is outside this claim.
Figure 1 reads left to right. Each row is one reclaim policy. The middle cells are a small compressed pool. The right side is the swap disk. Red is a page the next call will touch. Stock reclaim on a classic LRU mixes hot and cold pages into the pool and onto the disk, so the next call has to pick the red pages off the disk. A perfect prefetch on that layout reads the right pages, and the count and the locations do not change. The gain is only parallelism. The mechanism under test sends cold pages away first and stores the hot set last, so the hot set stays in the pool. The bottom row is what the later numbers say. Under MGLRU and zsmalloc, stock reclaim evicts by generation. Pages just used sit in the youngest generation and leave last, and most of them land in the pool.
TL;DR
A long-lived sandbox is reclaimed into zswap and swap, and the next call has to read those pages back. The natural guess is that stock reclaim scatters the next call’s pages across the disk, that the set is stable across calls, and that storing it into the compressed pool last will beat any prefetch that only runs on the restore side.
On 6.1 that mechanism is 0.46 to 0.55 of the strongest stock, and 0.61 to 0.64 of a perfect restore-side prefetch. On 6.12 it is only 0.87 to 0.94 of the perfect prefetch, and repeats swing from 0.73 to 1.19. On the same disk, stock reads 42.3 MB per wakeup on 6.1 and 13.7 MB on default 6.12. Turning MGLRU off returns 35.5 MB. MGLRU with zbud is 27.7 MB.
The room for a placement change is the disk stock still reads in the same setting. MGLRU plus a dense zpool has already taken most of it.
The hypothesis under test
Stock reclaim does not use information from the previous call. After a full reclaim, the pages the next call reads back sit on swap in runs of about 3.5 pages. Reclaim order, ordered pageout, and a two-phase reclaim do not assemble them into one layout. That was measured on 6.1, and it holds.
The activation set is stable, so a reclaim can send away pages that were not in the previous set first, with the pool share at zero and writeback allowed, and only then store the activation set into the pool and forbid writeback. Wakeup then only decompresses those pages. That should beat a stock layout plus a perfect restore-side prefetch, because the prefetch can only read the same bytes earlier. It looked feasible on 6.1 and SATA. Stock read 40 to 90 MB per wakeup, and storing the set in the pool left 5 to 38 MB.
How it was measured
Each sandbox runs in a Firecracker microVM with 2 vCPUs and 1 GB of memory. The guest has its own cgroup v2, zswap with lzo and a pool of 5% of memory, and a swap disk. The swap disk is a file on a host SATA SSD. Before each wakeup the host drops the page cache of that file, so the read hits the disk. The two disks are a datacenter Intel SATA SSD and a Samsung 870 SATA SSD. The guest kernels are kernel.org 6.1.102 and 6.12.111, both with idle page tracking. The 6.1 configuration is classic LRU and zbud. The 6.12 configuration has MGLRU on and uses zsmalloc.
The sandboxes are a TypeScript language server doing one diagnostic and completion step, pyright doing one Python type-check step, and a headless Chromium rendering and driving a page. Each measurement is 6 rounds. A round is a warm call, a full reclaim, and a wakeup call. The reported value is the median of rounds 2 through 6, because round 1 has no previous activation set. Each configuration is repeated 3 or 4 times, and the p95 pools every round and every repeat. The denominator is the better of two stock arms in the same batch. One arm is a one-shot full reclaim. The other reclaims at the same pool share as the mechanism under test.
The perfect restore-side prefetch comes from a snapshot. Before a wakeup the whole VM is paused, a full snapshot is saved, and the swap disk is backed up. The wakeup then runs as stock, and the pages it actually reads back are recorded. The snapshot is restored, the same stock wakeup runs again, and that measures the restore overhead itself. One more restore issues a prefetch of the recorded set when the call arrives, and the layout is not touched. The recorded set covers 97% to 99.8% of the pages the prefetched run actually reads back. The reported prefetch time subtracts the restore overhead measured on that run, 0.1 to 0.3 seconds. Snapshot files and the backup sit on a memory disk. Restoring the swap disk writes back only the 4 KiB blocks that differ, so the measurement itself does not write to the disk under test. The first attempt did not do this. The snapshot path wrote several gigabytes to the same SATA disk, per-page reads rose from 99 us to 126 us, stock looked 53% slower, and that batch was discarded.
A pause-and-restore of the whole machine is a separate control. While idle, the guest is not reclaimed. The VM is paused, a full snapshot is written, and the VMM is killed. When the call arrives, the guest pages the VMM actually mapped after the previous call are prefetched in parallel, and then the snapshot is loaded.
The same mechanism on two kernels
The vertical axis of Figure 2 is reactivation time over the strongest stock in the same batch. Lower is better. On 6.1, the three sandboxes put the blue bars at 0.46 to 0.55 and the orange bars at 0.76 to 0.86. Blue is about 0.6 of orange. On 6.12, orange is 0.82, 0.95, and 0.98, and blue is 0.74, 0.82, and 0.92. The average gap is 6% to 13%. The black points spread. Four tsserver repeats, relative to the perfect prefetch, are 1.03, 0.93, 0.82, and 0.96. The browser repeats are 0.73, 0.99, 1.05, and 1.19.
A perfect prefetch changes when the reads happen, not how many bytes they are. Storing the set in the pool changes the byte count. On 6.1 a stock wakeup reads 39 to 72 MB, and a perfect prefetch reads the same amount, only in parallel. After the set is in the pool the read is 5 to 38 MB. The gap on the left is those tens of megabytes. On 6.12 stock already reads 14 to 50 MB, and the perfect prefetch reads 14 to 49 MB. The mechanism drops tsserver to 3 MB and pyright to 15 MB. Disk is only part of the wakeup. The rest is decompression, fault handling, and the call itself, so the time barely moves. On the browser the mechanism reads 15 MB and the perfect prefetch reads 18 MB.
There is also a memory accounting problem. On 6.12 the mechanism’s idle memory is 1.6 to 2.1 times stock: 61 against 36 MB, 150 against 70 MB, and 103 against 66 MB. Most of the extra is swap cache. When 6.12 zswap reads a page back it deletes the pool entry, and a prefetched page stays in the swap cache as a dirty page. With the pool share frozen and writeback forbidden, those pages can neither re-enter the pool nor go to disk. Forcing them onto disk brings memory in line with stock, and on 6.12 it makes the browser 1.45 times stock. At the same real memory, the 6.12 gap against a perfect prefetch only gets smaller.
Both kernels were compared as one sandbox, a 5% pool, and SATA. In an earlier measurement a 20% pool holds the whole sandbox, and no policy wins. On NVMe the 6.1 advantage retreats from about 0.4 to about 0.7, because each read is cheaper and the reads that were avoided buy less time.
What actually changed the disk reads
6.1 and 6.12 differ by more than two years of kernel work, and by two configuration choices. To separate them, the same Samsung 870 SATA disk and the same tsserver workload run only the one-shot stock reclaim, switching those two choices one at a time. On 6.12, MGLRU is turned off at runtime and the zpool is switched to zbud. On 6.1, the zpool is switched to zsmalloc. Each point is 4 repeats. This step does not rebuild the kernel.
From left to right, 6.1 with zbud reads 42.3 MB. 6.1 with zsmalloc reads 33.9 MB. On 6.12 with MGLRU off, zbud is 33.7 MB and zsmalloc is 35.5 MB, the same band as 6.1 with zsmalloc. MGLRU on and zbud is 27.7 MB. Both on is 13.7 MB.
Everything in 6.1 to 6.12 other than these two choices has no measurable effect on this workload. 33.9 MB sits next to 33.7 to 35.5 MB. MGLRU alone saves about 6 MB. Switching to zsmalloc alone does not save anything on 6.12. Together they save 22 MB, a factor of 2.6. A reading that matches the numbers is this. MGLRU makes a full reclaim walk generations in order. Pages the call just touched are in the youngest generation and are reclaimed last. If the pool still has room, they enter it. Whether the pool can hold those late arrivals depends on its density. A zbud pool is already full of colder pages reclaimed earlier. A zsmalloc pool still has room. MGLRU supplies the order that stores the hot set last. zsmalloc supplies the capacity that fits it. Those are the two things the mechanism under test does by hand.
The bisect uses only tsserver. Upstream 6.1 contains MGLRU, but this 6.1 configuration was built without it, so a 6.1 with MGLRU is inferred from the 6.12 switch and was not measured directly. Stock on 6.12 also moves a lot between cells. The tsserver median runs from 0.73 to 1.17 seconds, and the page count from 2.2k to 7.5k. The judgment here uses the disk bytes and the per-page cost across four repeats, not the time of one cell.
What a pause to disk buys
On 6.12, pause-and-restore of the whole machine is 1.24, 0.83, and 1.36 times the strongest stock on the three sandboxes. The p95 values are 0.94, 0.60, and 1.17 times. Idle host memory is zero. The in-guest call itself adds almost no time. The time sits in two places. Before wakeup, prefetching the working set reads 270 to 480 MB, which is 0.55 to 0.99 seconds. Each pause writes a 1 GB memory file, which is 3.7 to 4.0 seconds.
On idle memory alone, this is the cheapest. A sandbox called every T seconds writes 1 GB on every idle period, and write bandwidth becomes the capacity limit. A small model counts both constraints. During the call, the reclaim, or the pause, the sandbox holds hot memory. The rest of the time it holds idle memory. Bytes written per idle period, divided by the disk’s write bandwidth, and the memory budget, together cap how many sandboxes fit. Write bandwidth is the fastest rate measured during these pauses, 293 MB/s. The host gives the sandboxes 64 GiB. For T from 30 to 300 seconds, every policy hits the write bandwidth first. Reclaim writes 23 to 77 MB per idle period. The snapshot writes 1 GB. The snapshot then fits 5 to 35 times fewer sandboxes. Even a diff snapshot that writes only the working set, 270 to 480 MB, which is an unmeasured lower bound, still fits 4.6 to 9.4 times fewer. This depends on the pause and the swap sharing one disk. If snapshots go to another fast disk, or the call interval stretches to hours, memory is the bottleneck again, and the snapshot’s zero idle memory wins.
Why the conjecture does not hold on 6.12
Figure 2 on the right and Figure 3 together decide it. On default 6.12, a perfect restore-side prefetch is already close to the mechanism. The prefetch did not get stronger. Stock reclaim already keeps most of the hot set in the pool, and a wakeup has only 14 to 50 MB of disk left to avoid. The mechanism still works on 6.1, and at the same memory it is still about 40% ahead of a perfect prefetch. What does not hold is the premise it needs, that stock reclaim scatters the hot set onto the disk. MGLRU and zsmalloc together already cover most of that premise in the default configuration.
A few gaps remain. The 6.12 comparison is one sandbox and one pool size. Several sandboxes sharing a pool have no clean perfect-prefetch control, and the direction is unknown. The mechanism uses more memory on 6.12, which biases the comparison in its favor. After the extra pages are drained it is slower, so at the same memory the conclusion is worse for it. The perfect prefetch subtracts restore overhead, which makes it look optimistic. The subtracted amount is 0.1 to 0.3 seconds, the same order as the gap, and it does not change the judgment that there is no stable 20% win.
What still sets a wakeup’s disk reads
The ceiling of a placement change is the disk stock still reads in the same setting, times the cost of one read. Wakeup cost is roughly read count times one-read latency, plus a fixed part. A mechanism that only moves pages can save at most the disk term. It applies on a restore path after swap or zswap reclaim. The smaller the disk share, the lower the ceiling. Default 6.12 leaves 14 to 50 MB per wakeup.
On a classic LRU, a prefetch set that misses no page still saves only 10% to 25%. On 6.1 guests, SATA, a 5% pool, and three sandboxes, a perfect prefetch is 0.76 to 0.86 of the strongest stock, because it still reads 39 to 72 MB. Keeping the same pages in the compressed pool drops the read to 5 to 38 MB and the time to 0.46 to 0.55. Knowing what to read is not the same as being able to read it cheaply.
On 6.12, MGLRU and a dense zpool together cut stock wakeup reads by 2.6 times. Either one alone cuts them by at most about 1.3 times. On one Samsung 870 SATA disk, tsserver, zswap at 5%, and four repeats, reads fall from 35.5 MB with MGLRU off to 13.7 MB with MGLRU and zsmalloc. MGLRU alone is 27.7 MB. A reclaim-order or placement idea has to be measured against stock on the target kernel’s default LRU and default zpool. A gain on classic LRU does not carry over.
A policy that pauses to disk on every idle period runs out of write bandwidth before it runs out of memory. The inputs are measured. A snapshot writes 1 GB per idle period, reclaim writes 23 to 77 MB, and the disk writes at about 293 MB/s. This holds when the pause and the swap share one disk and the call interval is within minutes. It does not hold for a longer interval, or when snapshots go to their own fast disk.
Writes made by the measurement itself cannot land on the disk being read. After snapshots and backups wrote several gigabytes to the same SATA SSD, the later per-page read cost rose from 99 us to 126 us, stock looked 53% slower, and the direction of a comparison can flip. Moving those writes to memory, and writing back only the blocks that differ, brings the read cost back to 73 to 103 us.
These numbers cover one 1 GB microVM, a SATA disk, and a compressed pool around 5%. They do not cover 6.12 on NVMe, a clean comparison of several sandboxes sharing a pool, an open-loop arrival process, or a direct measurement of 6.1 with MGLRU compiled in.
The hot set already stays in the pool
On 6.12 the default reclaim already stores the hot set last, and the dense pool has room for it. Doing the same thing by hand no longer wins a stable 20% over a perfect prefetch. The disk a wakeup still reads is what is left to optimize, and on this kernel most of it is already gone.
Related analysis
Agent Sandbox Wakeup: What to Prefetch Matters More Than When asks whether the lead time before a tool call can hide the restore. Which pages come back matters more than when the fetch starts.
在经典 LRU 上,把下一次要用的页留在压缩池里,比知道它们在哪还值钱。换到 MGLRU 加 zsmalloc 的 6.12,stock 回收自己就做到了大半。唤醒读盘从 42 MB 降到 14 MB,剩下的空间不够任何放置再稳定地赢 20%。
给 agent 用的沙箱越来越像长期进程。一个语言服务器、一套 Python 工具,或者一个无头浏览器,在两次调用之间空闲几秒到几分钟,等模型,也等人。宿主想在一台机器上多放几个,就会在空闲时把内存回收掉,下一次调用再付唤醒的代价。这篇文章讲的是这一次唤醒花在哪,以及一个看起来自然的放置,为什么在较新的内核上几乎没有空间。读完应该能判断,一个针对回收后恢复的放置或预取,值不值得在你正在跑的内核上做。
回收之后,一次调用要读回什么
Linux 回收匿名页时,把页写到 swap 设备上,页表项里只留一个 swap 槽位号。下一次访问触发缺页,内核按槽位把页读回来。SATA SSD 上读一页 4 KiB 要几十到一百多微秒。唤醒代价大致就是要读回的页数乘上单次读的代价,再加上缺页处理本身。一个 tsserver 沙箱在全量回收之后,一次调用通常要读回 50 到 70 MB,一万多页。
zswap 在这条路径上加了一层。它是内核里的压缩 swap 缓存。页在写到 swap 设备之前,先压缩进内存里的池子。池子满了,再按 LRU 把最旧的条目写回真正的盘(Linux kernel docs, zswap)。池子大小由 max_pool_percent 限定,常见配置是内存的几个百分点。还在池里的页只要解压,几微秒。已经离开池子的页要读盘。池子装不下整个沙箱时,哪些页留在池里,就决定了唤醒要读多少盘。6.8 之后的内核还给每个 cgroup 加了 memory.zswap.max 和 memory.zswap.writeback,可以限定一个 cgroup 在池里的份额,也可以禁止它的条目被写回盘。内存有压力时,shrinker 还可以主动把冷条目写回盘。
池里的条目放在分配器 zpool 里。zbud 每个物理页最多放两个压缩对象,简单,密度低。zsmalloc 按大小分类打包,密度高得多。同样 5% 的池子,zsmalloc 能放下的页明显更多。6.1 时代很多配置默认 zbud。6.12 的默认是 zsmalloc。
另一个变量是回收按什么顺序挑页。经典 LRU 只有 active 和 inactive 两条链表。一次 memory.reclaim 把整个 cgroup 收到接近零,挑页的先后和下一次调用会不会用到这页关系不大(Linux cgroup v2)。多代 LRU,也就是 MGLRU,按年龄把页分成若干代,从最老的一代开始回收。页被访问之后,在老化时会被提升到最年轻的一代。最年轻的两代相当于 active 链表,不会被先淘汰(Multi-Gen LRU,设计文档)。MGLRU 在 6.1 合入主线。6.12 的很多发行版配置默认打开它。
预取和整机快照已经有了
已有的一种做法是恢复侧预取。调用到达时,在缺页之前,把预计要用的页提前发出读请求,比如对这些地址做 MADV_WILLNEED。它不改页在哪。它只是把串行的缺页变成并行的读。上界是预取的集合正好等于这次调用要用的集合。这时剩下的代价,是把这些页从它们原来所在的位置读回来。
另一种做法是整机快照。Firecracker 是面向 serverless 的轻量虚拟机监视器,AWS Lambda 和 Fargate 在用(Agache et al. 2020, NSDI)。快照由客户机内存文件和 microVM 状态文件组成,可以是完整拷贝,也可以只含脏页。恢复时,内存文件可以交给内核按缺页读入,也可以交给用户态的 userfaultfd 处理程序(Firecracker snapshot support)。按缺页读入时,REAP 测到从快照启动的执行时间平均比常驻内存高 95%。同一个函数跨调用会访问一个稳定的工作集,于是 REAP 在第一次调用时记下这个集合,写成一个紧凑文件,之后恢复时一次读进来,冷启动降了 3.7 倍(Ustiugov et al. 2021, ASPLOS)。快照的好处是空闲时宿主上一字节内存都不占。代价是每次暂停都要写出内存状态。
两种做法都依赖同一件事。跨调用的工作集是稳定的。在这些沙箱上,这一点成立。tsserver 相邻两次调用的激活集重合 0.82 到 0.84。Python 工具集是 0.94。一个 torch 推理沙箱是 0.996。随机访问的 JVM 负载只有 0.11,不在这个判断里面。
图 1 从左往右读。每一行是一种回收策略。中间两格是容量很小的压缩池,右边是 swap 盘,红色是下一次调用会碰的页。第一行是经典 LRU 的 stock。热页和冷页混着进池,也混着落盘,下一次调用要去盘上把分散的红页读回来。第二行假设预取是完美的。读的对象一页不差,读的位置和数量没有变,改进只来自并行。第三行是被检验的机制。先把冷页送走,最后才存热集,热集于是留在池里。第四行是后面的数据要说明的事。在 MGLRU 和 zsmalloc 下,stock 按代淘汰。刚刚用过的页在最年轻一代,最后离开,大部分也落进了池里。
TL;DR
长期沙箱被回收到 zswap 和 swap 之后,下一次调用要把刚才用过的页读回来。一个自然的猜想是,stock 回收会把这些页打散到盘上,而这组页跨调用是稳定的,所以回收时把它最后存进压缩池,会比任何只在恢复侧做的预取都好。
在 6.1 上,这个机制是最强 stock 的 0.46 到 0.55 倍,是完美恢复侧预取的 0.61 到 0.64 倍。在 6.12 上,它只剩完美预取的 0.87 到 0.94,重复之间从 0.73 摆到 1.19。同一块盘上,stock 每次唤醒的读盘量在 6.1 是 42.3 MB,6.12 默认是 13.7 MB。关掉 MGLRU 回到 35.5 MB。只开 MGLRU、仍用 zbud,是 27.7 MB。
放置还能动的空间,等于 stock 在同一场景下还要读的盘。MGLRU 加上高密度的 zpool,已经把其中大部分拿走了。
被检验的假设
stock 回收不用上一轮调用的信息。一次全量回收之后,下一次调用要读回的页在 swap 上碎成大约每段 3.5 页。回收顺序、有序 pageout、分成两段的回收,这些 stock 原语都组不出把它们放在一起的布局。这一点在 6.1 上测过,成立。
激活集是稳定的,所以可以在回收时先把不在上一轮激活集里的页送走。这时压缩池的份额为零,允许写回盘。最后才把激活集存进池里,并禁止它被写回。唤醒时这组页只需解压。这应该比 stock 布局加上完美的恢复侧预取更好,因为后者只能更早地读同样多的盘。在 6.1 和 SATA 上,这件事当时看起来可行。stock 每次唤醒要读 40 到 90 MB。把集合放进池之后,只读 5 到 38 MB。
怎么测的
每个沙箱跑在一个 Firecracker microVM 里,2 个 vCPU,1 GB 内存。客户机里有自己的 cgroup v2、zswap(lzo,池子是内存的 5%)和一块 swap 盘。swap 盘是宿主上一块 SATA SSD 的文件。每次唤醒之前,宿主丢掉这个文件的页缓存,保证读真的落到盘上。两块盘分别是一块数据中心级 Intel SATA SSD,和一块 Samsung 870 SATA SSD。客户机内核是 kernel.org 的 6.1.102 和 6.12.111,都打开了 idle page tracking。6.1 的配置是经典 LRU 和 zbud。6.12 打开 MGLRU,zpool 用 zsmalloc。
沙箱有三种。TypeScript 语言服务器做一步代码诊断和补全。pyright 做一步 Python 类型检查。无头 Chromium 渲染并操作一个页面。每次测量跑 6 轮。一轮是一次热调用、一次全量回收、一次唤醒调用。报告第 2 到第 6 轮的中位数,因为第 1 轮没有上一轮的激活集。每个配置重复 3 到 4 次。p95 把所有轮和所有重复合在一起算。分母是同一批次里两种 stock 里更好的那一个。一种是一次性全量回收。另一种按与被检验机制相同的池份额回收。
完美的恢复侧预取用快照得到。每次唤醒之前,把整个 VM 暂停,存一个完整快照,并备份 swap 盘。先按 stock 正常跑这一次唤醒,记下它实际读回的页。然后恢复快照,再跑一遍同样的 stock 唤醒,量出恢复本身的开销。再恢复一次。这一次在调用到达时,对刚才记下的真集合发出预取,布局一点不动。记下的集合覆盖了预取那一遍实际读回页的 97% 到 99.8%。报告的完美预取时间减去了那一遍量出的恢复开销,0.1 到 0.3 秒。快照文件和备份放在内存盘上。恢复 swap 盘时只写回不同的 4 KiB 块,所以测量本身不往被测的盘写东西。第一次没有这样做。快照路径往同一块 SATA 写了数 GB,之后每页读盘从 99 us 涨到 126 us,stock 看起来慢了 53%,那一批作废。
整机快照的暂停恢复是另一个对照。空闲时不在客户机里回收,直接把 VM 暂停,写完整快照,杀掉 VMM。调用到达时,先按上一次调用后 VMM 实际映射的客户机页做并行预取,再加载快照。
两个内核上的结果
6.1 的三类沙箱里,蓝色在 0.46 到 0.55,橙色在 0.76 到 0.86,蓝色大约是橙色的 0.6。6.12 上,橙色是 0.82、0.95、0.98,蓝色是 0.74、0.82、0.92,平均只比橙色低 6% 到 13%。黑点散得很开。tsserver 的四次重复相对完美预取是 1.03、0.93、0.82、0.96。browser 是 0.73、0.99、1.05、1.19。
完美预取改变的是读的时机,不是读的量。放进池改变的是读的量。在 6.1 上,stock 唤醒要读 39 到 72 MB,完美预取也要读这么多,只是并行发出。放进池之后只读 5 到 38 MB。左图的差距基本就是这几十 MB。在 6.12 上,stock 本来就只读 14 到 50 MB,完美预取读 14 到 49 MB。被检验的机制在 tsserver 上降到 3 MB,在 pyright 上降到 15 MB。读盘只占唤醒时间的一部分,剩下的是解压、缺页处理和调用本身,所以时间只省了一点。browser 上,这个机制读 15 MB,完美预取读 18 MB,几乎没有差别。
这里还有一个内存口径的问题。在 6.12 上,这个机制空闲时占的内存是 stock 的 1.6 到 2.1 倍,三组分别是 61 对 36 MB、150 对 70 MB、103 对 66 MB。多出来的主要是 swap cache。6.12 的 zswap 读回一页时会把条目从池里删掉,预取读回的页以脏页的形式留在 swap cache 里。池份额被冻住、又禁止写回时,这些页既不能再进池,也不能写到盘上。强行把它们送到盘上,内存可以和 stock 持平,但在 6.12 上这让 browser 变成 stock 的 1.45 倍。所以在同样的真实内存下,6.12 上和完美预取的差距只会更小。
两个内核上的比较都是单沙箱、池子 5%、SATA。更早的测量里,池子放大到 20%,整个沙箱都装得下,任何策略都没有收益。换到 NVMe,6.1 上的优势也从大约 0.4 退到大约 0.7,因为每次读更便宜,省下的读次数值的时间更少。
读盘量是被什么改变的
6.1 和 6.12 之间差了两年多的内核演进,也差了两个配置选择。为了把它们分开,我们在同一块 Samsung 870 SATA、同一个 tsserver 负载上,只跑 stock 的一次性回收,逐个切换这两个选择。6.12 上运行时关掉 MGLRU,并把 zpool 换成 zbud。6.1 上把 zpool 换成 zsmalloc。每个点 4 次重复。这一步不需要重新编译内核。
从左往右,6.1 用 zbud 是 42.3 MB。6.1 换 zsmalloc 是 33.9 MB。6.12 关掉 MGLRU 时,zbud 是 33.7 MB,zsmalloc 是 35.5 MB,和 6.1 的 zsmalloc 在同一水平。6.12 打开 MGLRU、仍用 zbud,是 27.7 MB。两者都开,是 13.7 MB。
先看直接的结论。6.1 到 6.12 之间,除了这两个选择,其余代码在这个负载上没有可测的影响。33.9 MB 旁边就是 33.7 到 35.5 MB。只开 MGLRU,大约少 6 MB。在 6.12 上只换 zsmalloc,不降。两者一起少 22 MB,是 2.6 倍。和这些数字一致的一种读法是这样。MGLRU 让全量回收按代的顺序走。刚刚被调用碰过的页在最年轻的一代,最后才被回收。这时池子里还有空位,它们就进了池。池子能不能装下这些最后到来的页,取决于密度。zbud 的池子在这之前已经被更早回收的冷页填满。zsmalloc 的池子还有余量。MGLRU 提供了热集最后存的顺序。zsmalloc 提供了让它们装得下的容量。这正是被检验的机制用显式手段做的两件事。
这次 bisect 只用了 tsserver 一种负载。6.1 主线里有 MGLRU,但这套 6.1 配置没有编译它。所以带 MGLRU 的 6.1 是不是也这样,是从 6.12 上的开关推出来的,没有直接测。6.12 的 stock 在格子之间波动也大。tsserver 每格的中位数在 0.73 到 1.17 秒之间,读的页数在 2.2k 到 7.5k 之间。这里的判断靠的是四次重复的读盘量和每页读成本,不是某一格的时间。
暂停到盘换来的是什么
整机快照的暂停恢复,在 6.12 上的端到端唤醒时间是最强 stock 的 1.24、0.83、1.36 倍。p95 是 0.94、0.60、1.17 倍。空闲时宿主内存为 0。客户机里的调用本身几乎不多花时间。时间花在两处。唤醒前按工作集预取 270 到 480 MB,是 0.55 到 0.99 秒。每次暂停写 1 GB 的内存文件,是 3.7 到 4.0 秒。
只看空闲内存,它是最省的。一个沙箱每隔 T 秒被调用一次,每次空闲都要写 1 GB,写带宽就成了容量的上限。一个简单模型把两者一起算。沙箱在调用、回收或暂停期间占着热内存,其余时间占空闲内存。每次空闲写出的字节数除以盘的写带宽,再和内存预算一起,决定能放多少沙箱。写带宽取这些暂停里测到的最快值,293 MB/s。宿主给沙箱 64 GiB。T 在 30 到 300 秒时,所有策略都先被写带宽卡住。回收类策略每次空闲写 23 到 77 MB。快照写 1 GB。于是快照能放的沙箱数少 5 到 35 倍。就算假设差异快照只写工作集,270 到 480 MB,这是没测过的下界,仍然少 4.6 到 9.4 倍。这个结论依赖暂停和 swap 共用一块盘。如果快照写到另一块很快的盘上,或者调用间隔长到小时级,内存重新成为瓶颈,快照的空闲零内存就会占优。
为什么这个判断在 6.12 上不成立
图 2 的右边和图 3 合在一起就是决定性的证据。在 6.12 的默认配置下,完美的恢复侧预取已经和这个机制差不多。原因不是预取变强了。stock 回收本身已经把大部分热集留在了池里,每次唤醒只剩 14 到 50 MB 的读盘可省。机制本身没有失效。它在 6.1 上仍然有效,同内存下仍比完美预取好大约 40%。不成立的是它依赖的前提,也就是 stock 回收会把热集打散到盘上。MGLRU 和 zsmalloc 的组合,已经覆盖了这个前提在默认配置下的大部分。
还有几处没测到。6.12 上的比较只有单沙箱和一种池子大小。多沙箱共享池的情形,没有干净的完美预取对照,方向不明。这个机制在 6.12 上多占内存,对比偏向它。把多出来的页排空之后反而更慢,所以按同样的内存算,结论对它更不利。完美预取扣除了恢复开销,这使它偏乐观。扣掉的量是 0.1 到 0.3 秒,和差距同一量级,不改变没有稳定赢 20% 这个判断。
唤醒读盘还由什么决定
放置的上界,是 stock 在同一场景下还要读的盘,乘上单次读的延迟。唤醒代价近似为读盘次数乘单次延迟,再加上一个固定部分。任何只改变页放在哪的机制,最多省掉读盘那一项。它在 swap 或 zswap 回收之后的恢复路径上成立。读盘占比越小,上界越低。6.12 的默认配置下,每次唤醒只剩 14 到 50 MB。
在经典 LRU 上,即使预取的集合一页不差,恢复侧预取也只省 10% 到 25%。6.1 客户机、SATA、池子 5%、三类沙箱上,完美预取是最强 stock 的 0.76 到 0.86,因为它仍要读 39 到 72 MB。把同一组页留在压缩池里,读盘降到 5 到 38 MB,时间降到 0.46 到 0.55。知道要读什么,不等于能便宜地读回来。
在 6.12 上,MGLRU 和高密度 zpool 一起,把 stock 唤醒的读盘降了 2.6 倍。两者单独,最多大约降 1.3 倍。同一块 Samsung 870 SATA、tsserver、zswap 5%、四次重复,读盘从 MGLRU 关闭时的 35.5 MB 降到 MGLRU 加 zsmalloc 的 13.7 MB。只开 MGLRU 是 27.7 MB。任何回收顺序或放置的想法,都要在目标内核的默认 LRU 和默认 zpool 下测 stock。经典 LRU 上的收益不能外推。
每次空闲都暂停到盘的策略,容量先被写带宽卡住,而不是被内存卡住。输入是实测的。快照每次空闲写 1 GB,回收写 23 到 77 MB,盘的写带宽大约 293 MB/s。条件是暂停写和 swap 读共用一块盘,调用间隔在分钟以内。间隔更长,或者快照写到自己的快盘上,这个判断不成立。
测量机制自己的写入,不能落在被测的读盘设备上。快照和备份往同一块 SATA SSD 写了数 GB 之后,后续测量的每页读成本从 99 us 涨到 126 us,stock 看起来慢了 53%,对比的方向也会翻。把这些写入移到内存,只写回有差异的块,读成本回到 73 到 103 us。
这些数字只覆盖单个 1 GB 的 microVM、SATA 盘,以及 5% 左右的压缩池。没有覆盖 NVMe 上的 6.12,没有多沙箱共享池的干净对比,没有开放环的真实到达过程,也没有直接测编译进 MGLRU 的 6.1。
热集已经留在池里
6.12 的默认回收已经把热集留到最后,高密度的池也装得下。再用显式手段做同一件事,相对完美预取已经没有稳定的 20% 可赢。唤醒还要读的那部分盘,才是还能优化的对象。在这个内核上,其中大部分已经不在了。
相关分析
Agent 沙箱唤醒:预取什么比何时预取更重要 问的是工具调用之前的提前量能不能把恢复藏住。取回哪些页,比何时开始取更重要。