Agent Sandbox: What Can Still Cut Runtime Memory and Page Faults

On a classic LRU, keeping the next call's pages in the compressed pool beats knowing where they are. On Linux 6.12, with MGLRU and zsmalloc, stock reclaim already does most of that. Wakeup disk reads fall from 42 MB to 14 MB, and what remains is not enough for a placement change to win another 20%.

A sandbox for an agent is starting to look like a long-lived process. A language server, a Python tool set, or a headless browser sits idle for seconds to minutes between calls, waiting on the model or on a person. The host wants more of them on one machine, so it reclaims their memory while they wait, and the next call pays to wake them up. This article is about where that wakeup spends its time, and why a placement idea that looks natural has almost no room left on a newer kernel. After reading it you should be able to tell whether a placement or prefetch idea for post-reclaim restore is worth building on the kernel you actually run.

After reclaim, a call has to read pages back

When Linux reclaims an anonymous page, it writes the page to a swap device and leaves a swap slot number in the page table. The next access faults, and the kernel reads the page back by that slot. A 4 KiB read on a SATA SSD is tens to more than a hundred microseconds, so the wakeup cost is roughly the number of pages read back, times the cost of one read, plus the fault handling itself. After a full reclaim, one call into a tsserver sandbox usually reads 50 to 70 MB back, a bit over ten thousand pages.

zswap adds a layer on that path. It is a compressed swap cache in the kernel. A page is compressed into an in-memory pool before it is written to the swap device, and only when the pool is full does the oldest entry get written out, in LRU order (Linux kernel documentation, zswap). The pool size is capped by max_pool_percent, often a few percent of memory. A page that is still in the pool costs a decompression of a few microseconds. A page that has left costs a disk read. When the pool cannot hold the whole sandbox, which pages stay in it decides how much disk the wakeup reads. Kernels since 6.8 also give each cgroup memory.zswap.max and memory.zswap.writeback, so a cgroup can be capped in the pool and can be forbidden from writeback, and a shrinker can push cold entries to disk under memory pressure.

Entries in the pool live in an allocator, the zpool. zbud stores at most two compressed objects in one physical page. It is simple, and the density is low. zsmalloc packs objects by size class, and the density is much higher. The same 5% pool holds many more pages with zsmalloc. A lot of 6.1 configurations defaulted to zbud. The 6.12 default is zsmalloc.

The other variable is the order in which reclaim picks pages. Classic LRU has two lists, active and inactive. A memory.reclaim that drains a whole cgroup toward zero does not pick pages according to whether the next call will use them (Linux cgroup v2 documentation). Multi-gen LRU, MGLRU, sorts pages into generations by age and reclaims from the oldest generation. A page that is accessed is promoted to the youngest generation when generations age. The two youngest generations are the equivalent of the active list, and they are not evicted first (Linux kernel documentation, Multi-Gen LRU, design notes). MGLRU landed in 6.1. Many distribution configurations turn it on by default in 6.12.

Prefetch and a full snapshot already exist

One existing move prefetches on the restore side. When the call arrives, and before the faults, it issues reads for the pages it expects to need, for example with MADV_WILLNEED. It does not move the pages. It turns serial faults into parallel reads. Its ceiling is a prefetch set that matches the set this call will touch. What remains is the cost of reading those pages from wherever they already sit.

The other move is a whole-machine snapshot. Firecracker is a lightweight hypervisor built for serverless, used by AWS Lambda and Fargate (Agache et al. 2020, NSDI). A snapshot is a guest-memory file plus a microVM state file. It can be a full copy or a diff of dirty pages. On restore, the kernel can fault the memory file in, or a userfaultfd handler in userspace can serve it (Firecracker snapshot support). When pages come in on fault, REAP measured snapshot start as 95% slower than a resident start, on average. The same function touches a stable working set across calls, so REAP records that set on the first call, writes it as one compact file, and reads it back in one shot on later restores. Cold start drops by 3.7 times (Ustiugov et al. 2021, ASPLOS). A snapshot costs the host no memory while the sandbox is idle. Every pause has to write the memory state out.

Both moves rely on a stable working set across calls. On these sandboxes that holds. Adjacent calls overlap by 0.82 to 0.84 for tsserver, 0.94 for the Python tool set, and 0.996 for a torch inference sandbox. A JVM workload with random access overlaps by 0.11, and it is outside this claim.

Four reclaim policies and where the next call's pages land. Classic LRU scatters them onto the swap disk. A perfect prefetch reads the same pages earlier. Storing the hot set last keeps it in the compressed pool. Stock reclaim with MGLRU and a dense zpool already leaves the youngest generation in the pool.
Figure 1. Where the pages the next call will touch (red) sit after reclaim. Classic LRU mixes them onto the disk. A perfect prefetch reads the same pages, only earlier and in parallel. Storing the hot set last keeps it in the compressed pool. With MGLRU and a dense zpool, stock reclaim already does that.

Figure 1 reads left to right. Each row is one reclaim policy. The middle cells are a small compressed pool. The right side is the swap disk. Red is a page the next call will touch. Stock reclaim on a classic LRU mixes hot and cold pages into the pool and onto the disk, so the next call has to pick the red pages off the disk. A perfect prefetch on that layout reads the right pages, and the count and the locations do not change. The gain is only parallelism. The mechanism under test sends cold pages away first and stores the hot set last, so the hot set stays in the pool. The bottom row is what the later numbers say. Under MGLRU and zsmalloc, stock reclaim evicts by generation. Pages just used sit in the youngest generation and leave last, and most of them land in the pool.

TL;DR

A long-lived sandbox is reclaimed into zswap and swap, and the next call has to read those pages back. The natural guess is that stock reclaim scatters the next call’s pages across the disk, that the set is stable across calls, and that storing it into the compressed pool last will beat any prefetch that only runs on the restore side.

On 6.1 that mechanism is 0.46 to 0.55 of the strongest stock, and 0.61 to 0.64 of a perfect restore-side prefetch. On 6.12 it is only 0.87 to 0.94 of the perfect prefetch, and repeats swing from 0.73 to 1.19. On the same disk, stock reads 42.3 MB per wakeup on 6.1 and 13.7 MB on default 6.12. Turning MGLRU off returns 35.5 MB. MGLRU with zbud is 27.7 MB.

The room for a placement change is the disk stock still reads in the same setting. MGLRU plus a dense zpool has already taken most of it.

The hypothesis under test

Stock reclaim does not use information from the previous call. After a full reclaim, the pages the next call reads back sit on swap in runs of about 3.5 pages. Reclaim order, ordered pageout, and a two-phase reclaim do not assemble them into one layout. That was measured on 6.1, and it holds.

The activation set is stable, so a reclaim can send away pages that were not in the previous set first, with the pool share at zero and writeback allowed, and only then store the activation set into the pool and forbid writeback. Wakeup then only decompresses those pages. That should beat a stock layout plus a perfect restore-side prefetch, because the prefetch can only read the same bytes earlier. It looked feasible on 6.1 and SATA. Stock read 40 to 90 MB per wakeup, and storing the set in the pool left 5 to 38 MB.

How it was measured

Each sandbox runs in a Firecracker microVM with 2 vCPUs and 1 GB of memory. The guest has its own cgroup v2, zswap with lzo and a pool of 5% of memory, and a swap disk. The swap disk is a file on a host SATA SSD. Before each wakeup the host drops the page cache of that file, so the read hits the disk. The two disks are a datacenter Intel SATA SSD and a Samsung 870 SATA SSD. The guest kernels are kernel.org 6.1.102 and 6.12.111, both with idle page tracking. The 6.1 configuration is classic LRU and zbud. The 6.12 configuration has MGLRU on and uses zsmalloc.

The sandboxes are a TypeScript language server doing one diagnostic and completion step, pyright doing one Python type-check step, and a headless Chromium rendering and driving a page. Each measurement is 6 rounds. A round is a warm call, a full reclaim, and a wakeup call. The reported value is the median of rounds 2 through 6, because round 1 has no previous activation set. Each configuration is repeated 3 or 4 times, and the p95 pools every round and every repeat. The denominator is the better of two stock arms in the same batch. One arm is a one-shot full reclaim. The other reclaims at the same pool share as the mechanism under test.

The perfect restore-side prefetch comes from a snapshot. Before a wakeup the whole VM is paused, a full snapshot is saved, and the swap disk is backed up. The wakeup then runs as stock, and the pages it actually reads back are recorded. The snapshot is restored, the same stock wakeup runs again, and that measures the restore overhead itself. One more restore issues a prefetch of the recorded set when the call arrives, and the layout is not touched. The recorded set covers 97% to 99.8% of the pages the prefetched run actually reads back. The reported prefetch time subtracts the restore overhead measured on that run, 0.1 to 0.3 seconds. Snapshot files and the backup sit on a memory disk. Restoring the swap disk writes back only the 4 KiB blocks that differ, so the measurement itself does not write to the disk under test. The first attempt did not do this. The snapshot path wrote several gigabytes to the same SATA disk, per-page reads rose from 99 us to 126 us, stock looked 53% slower, and that batch was discarded.

A pause-and-restore of the whole machine is a separate control. While idle, the guest is not reclaimed. The VM is paused, a full snapshot is written, and the VMM is killed. When the call arrives, the guest pages the VMM actually mapped after the previous call are prefetched in parallel, and then the snapshot is loaded.

The same mechanism on two kernels

Reactivation time divided by the best stock in the same batch. On guest kernel 6.1, storing the hot set last is well below stock plus a perfect prefetch. On 6.12 the two are close, and repeats of the hot-set policy cross the prefetch bars.
Figure 2. Reactivation time divided by the best stock in the same batch. The dashed line is 1. Orange is stock plus a perfect restore-side prefetch. Blue is the hot set stored last. Black points are repeats of the blue bar. On 6.1 the blue bars sit well below orange. On 6.12 they are close, and the points cross.

The vertical axis of Figure 2 is reactivation time over the strongest stock in the same batch. Lower is better. On 6.1, the three sandboxes put the blue bars at 0.46 to 0.55 and the orange bars at 0.76 to 0.86. Blue is about 0.6 of orange. On 6.12, orange is 0.82, 0.95, and 0.98, and blue is 0.74, 0.82, and 0.92. The average gap is 6% to 13%. The black points spread. Four tsserver repeats, relative to the perfect prefetch, are 1.03, 0.93, 0.82, and 0.96. The browser repeats are 0.73, 0.99, 1.05, and 1.19.

A perfect prefetch changes when the reads happen, not how many bytes they are. Storing the set in the pool changes the byte count. On 6.1 a stock wakeup reads 39 to 72 MB, and a perfect prefetch reads the same amount, only in parallel. After the set is in the pool the read is 5 to 38 MB. The gap on the left is those tens of megabytes. On 6.12 stock already reads 14 to 50 MB, and the perfect prefetch reads 14 to 49 MB. The mechanism drops tsserver to 3 MB and pyright to 15 MB. Disk is only part of the wakeup. The rest is decompression, fault handling, and the call itself, so the time barely moves. On the browser the mechanism reads 15 MB and the perfect prefetch reads 18 MB.

There is also a memory accounting problem. On 6.12 the mechanism’s idle memory is 1.6 to 2.1 times stock: 61 against 36 MB, 150 against 70 MB, and 103 against 66 MB. Most of the extra is swap cache. When 6.12 zswap reads a page back it deletes the pool entry, and a prefetched page stays in the swap cache as a dirty page. With the pool share frozen and writeback forbidden, those pages can neither re-enter the pool nor go to disk. Forcing them onto disk brings memory in line with stock, and on 6.12 it makes the browser 1.45 times stock. At the same real memory, the 6.12 gap against a perfect prefetch only gets smaller.

Both kernels were compared as one sandbox, a 5% pool, and SATA. In an earlier measurement a 20% pool holds the whole sandbox, and no policy wins. On NVMe the 6.1 advantage retreats from about 0.4 to about 0.7, because each read is cheaper and the reads that were avoided buy less time.

What actually changed the disk reads

6.1 and 6.12 differ by more than two years of kernel work, and by two configuration choices. To separate them, the same Samsung 870 SATA disk and the same tsserver workload run only the one-shot stock reclaim, switching those two choices one at a time. On 6.12, MGLRU is turned off at runtime and the zpool is switched to zbud. On 6.1, the zpool is switched to zsmalloc. Each point is 4 repeats. This step does not rebuild the kernel.

Swap-disk megabytes read per reactivation for stock reclaim on one disk and one workload. Reads stay near 34 to 42 MB until MGLRU and zsmalloc are both on, which drops the read to 13.7 MB.
Figure 3. Megabytes read from the swap disk per wakeup, stock reclaim only, one disk and one workload. Gray is MGLRU off, or 6.1 built without it. Blue is MGLRU on. Black points are four repeats. Both switches together land at 13.7 MB. Either one alone stays at 27.7 MB or above.

From left to right, 6.1 with zbud reads 42.3 MB. 6.1 with zsmalloc reads 33.9 MB. On 6.12 with MGLRU off, zbud is 33.7 MB and zsmalloc is 35.5 MB, the same band as 6.1 with zsmalloc. MGLRU on and zbud is 27.7 MB. Both on is 13.7 MB.

Everything in 6.1 to 6.12 other than these two choices has no measurable effect on this workload. 33.9 MB sits next to 33.7 to 35.5 MB. MGLRU alone saves about 6 MB. Switching to zsmalloc alone does not save anything on 6.12. Together they save 22 MB, a factor of 2.6. A reading that matches the numbers is this. MGLRU makes a full reclaim walk generations in order. Pages the call just touched are in the youngest generation and are reclaimed last. If the pool still has room, they enter it. Whether the pool can hold those late arrivals depends on its density. A zbud pool is already full of colder pages reclaimed earlier. A zsmalloc pool still has room. MGLRU supplies the order that stores the hot set last. zsmalloc supplies the capacity that fits it. Those are the two things the mechanism under test does by hand.

The bisect uses only tsserver. Upstream 6.1 contains MGLRU, but this 6.1 configuration was built without it, so a 6.1 with MGLRU is inferred from the 6.12 switch and was not measured directly. Stock on 6.12 also moves a lot between cells. The tsserver median runs from 0.73 to 1.17 seconds, and the page count from 2.2k to 7.5k. The judgment here uses the disk bytes and the per-page cost across four repeats, not the time of one cell.

What a pause to disk buys

On 6.12, pause-and-restore of the whole machine is 1.24, 0.83, and 1.36 times the strongest stock on the three sandboxes. The p95 values are 0.94, 0.60, and 1.17 times. Idle host memory is zero. The in-guest call itself adds almost no time. The time sits in two places. Before wakeup, prefetching the working set reads 270 to 480 MB, which is 0.55 to 0.99 seconds. Each pause writes a 1 GB memory file, which is 3.7 to 4.0 seconds.

On idle memory alone, this is the cheapest. A sandbox called every T seconds writes 1 GB on every idle period, and write bandwidth becomes the capacity limit. A small model counts both constraints. During the call, the reclaim, or the pause, the sandbox holds hot memory. The rest of the time it holds idle memory. Bytes written per idle period, divided by the disk’s write bandwidth, and the memory budget, together cap how many sandboxes fit. Write bandwidth is the fastest rate measured during these pauses, 293 MB/s. The host gives the sandboxes 64 GiB. For T from 30 to 300 seconds, every policy hits the write bandwidth first. Reclaim writes 23 to 77 MB per idle period. The snapshot writes 1 GB. The snapshot then fits 5 to 35 times fewer sandboxes. Even a diff snapshot that writes only the working set, 270 to 480 MB, which is an unmeasured lower bound, still fits 4.6 to 9.4 times fewer. This depends on the pause and the swap sharing one disk. If snapshots go to another fast disk, or the call interval stretches to hours, memory is the bottleneck again, and the snapshot’s zero idle memory wins.

Why the conjecture does not hold on 6.12

Figure 2 on the right and Figure 3 together decide it. On default 6.12, a perfect restore-side prefetch is already close to the mechanism. The prefetch did not get stronger. Stock reclaim already keeps most of the hot set in the pool, and a wakeup has only 14 to 50 MB of disk left to avoid. The mechanism still works on 6.1, and at the same memory it is still about 40% ahead of a perfect prefetch. What does not hold is the premise it needs, that stock reclaim scatters the hot set onto the disk. MGLRU and zsmalloc together already cover most of that premise in the default configuration.

A few gaps remain. The 6.12 comparison is one sandbox and one pool size. Several sandboxes sharing a pool have no clean perfect-prefetch control, and the direction is unknown. The mechanism uses more memory on 6.12, which biases the comparison in its favor. After the extra pages are drained it is slower, so at the same memory the conclusion is worse for it. The perfect prefetch subtracts restore overhead, which makes it look optimistic. The subtracted amount is 0.1 to 0.3 seconds, the same order as the gap, and it does not change the judgment that there is no stable 20% win.

What still sets a wakeup’s disk reads

The ceiling of a placement change is the disk stock still reads in the same setting, times the cost of one read. Wakeup cost is roughly read count times one-read latency, plus a fixed part. A mechanism that only moves pages can save at most the disk term. It applies on a restore path after swap or zswap reclaim. The smaller the disk share, the lower the ceiling. Default 6.12 leaves 14 to 50 MB per wakeup.

On a classic LRU, a prefetch set that misses no page still saves only 10% to 25%. On 6.1 guests, SATA, a 5% pool, and three sandboxes, a perfect prefetch is 0.76 to 0.86 of the strongest stock, because it still reads 39 to 72 MB. Keeping the same pages in the compressed pool drops the read to 5 to 38 MB and the time to 0.46 to 0.55. Knowing what to read is not the same as being able to read it cheaply.

On 6.12, MGLRU and a dense zpool together cut stock wakeup reads by 2.6 times. Either one alone cuts them by at most about 1.3 times. On one Samsung 870 SATA disk, tsserver, zswap at 5%, and four repeats, reads fall from 35.5 MB with MGLRU off to 13.7 MB with MGLRU and zsmalloc. MGLRU alone is 27.7 MB. A reclaim-order or placement idea has to be measured against stock on the target kernel’s default LRU and default zpool. A gain on classic LRU does not carry over.

A policy that pauses to disk on every idle period runs out of write bandwidth before it runs out of memory. The inputs are measured. A snapshot writes 1 GB per idle period, reclaim writes 23 to 77 MB, and the disk writes at about 293 MB/s. This holds when the pause and the swap share one disk and the call interval is within minutes. It does not hold for a longer interval, or when snapshots go to their own fast disk.

Writes made by the measurement itself cannot land on the disk being read. After snapshots and backups wrote several gigabytes to the same SATA SSD, the later per-page read cost rose from 99 us to 126 us, stock looked 53% slower, and the direction of a comparison can flip. Moving those writes to memory, and writing back only the blocks that differ, brings the read cost back to 73 to 103 us.

These numbers cover one 1 GB microVM, a SATA disk, and a compressed pool around 5%. They do not cover 6.12 on NVMe, a clean comparison of several sandboxes sharing a pool, an open-loop arrival process, or a direct measurement of 6.1 with MGLRU compiled in.

The hot set already stays in the pool

On 6.12 the default reclaim already stores the hot set last, and the dense pool has room for it. Doing the same thing by hand no longer wins a stable 20% over a perfect prefetch. The disk a wakeup still reads is what is left to optimize, and on this kernel most of it is already gone.

Agent Sandbox Wakeup: What to Prefetch Matters More Than When asks whether the lead time before a tool call can hide the restore. Which pages come back matters more than when the fetch starts.