Agent Sandbox Wakeup: What to Prefetch Matters More Than When

A tool call is visible in the token stream 1.7 seconds early, but the restore cost is which pages come back and how they are laid out on swap. In interactive sessions, reclaiming only the long gaps where a person is waiting already captures nine tenths of the memory benefit.

Coding agents, data-analysis agents, and agentic RL rollouts each give a session its own sandbox: an isolated environment that holds an interpreter, dependencies, a dataset, and browser processes. The sandbox spends most of its time waiting for the model, and the model spends most of its time waiting for a tool inside the sandbox to finish. When hundreds or thousands of sessions run together, the resident memory of idle sandboxes becomes the scarce resource, and swapping them out means the next tool call pays to wake them up. This article looks at a solution that looks clean. Before a tool call is fully generated, its name and arguments are already in the inference stream, so the sandbox can be restored during that lead time. We measured this on real sessions and real sandboxes. The lead time is real. It is almost never the quantity that decides the cost. What remains is an arithmetic condition for whether reactivation can be hidden, and three measured facts about how agent sandboxes use memory.

Why the sandbox waits, and why it is expensive

An agent session moves in rounds. The model reads context, writes a stretch of reasoning, and emits a tool call. The sandbox runs the tool and returns the result. The model starts the next round. Across 204 real coding-agent sessions, tool execution has a median of 0.16 seconds and a p90 of 7.1 seconds, while a model-inference stretch averages on the order of ten seconds. The idle gap between two tool calls splits at 600 seconds into two kinds. The short kind is waiting inside the tool loop, median 4.8 seconds. The long kind is waiting for a person to come back, or a paused session. That geometry is the space available for reclaim: a sandbox is idle almost all of the time.

The memory is not small. A data-science sandbox with pandas and a dataset resident holds several gigabytes of anonymous pages. A Chromium sandbox doing browser automation holds a bit over two gigabytes across nine processes. One characterization of agent systems, so far only on arXiv, reports a per-session sandbox working-set peak of 28 GB (AgentSysBench, arXiv 2608.15127). We did not reproduce that number. It is only a scale reference. Agentic RL makes the problem sharper. A group of rollouts is often 8 to 16 trajectories, each with a sandbox, and RollArt reports that environment-reset latency in production can reach hundreds of seconds in the tail (Gao et al. 2026, OSDI). Long-lived sandboxes also hold hot runtime state such as compiler daemons and language servers. Whether that state can be invalidated selectively is a separate question. This article is only about the memory pages.

Reclaim and restore: three existing lines

The first line swaps cold memory out transparently. Linux cgroup v2 offers memory.reclaim. Writing a byte count makes the kernel reclaim that much memory from the control group (Linux cgroup v2 documentation). The anonymous pages can go to two places. zram is a compressed in-memory block device. Pages stay in memory after compression, and swap-in is decompression (Linux zram documentation). A block-device swap writes pages to disk. Capacity is saved, and swap-in is slow. Google’s software-defined far memory uses zswap to compress cold pages proactively, holding about 20% of the cold data on average at an access cost of about 6 microseconds (Lagar-Cavilla et al. 2019, ASPLOS). Meta’s TMO uses pressure stall information, PSI, to measure the work lost to resource shortage and decides how much to offload from that, reclaiming every 6 seconds without application cooperation (Weiner et al. 2022, ASPLOS). This line solves capacity. Fetching a page back is a fault, which means restore starts after the access.

The second line prefetches from a recorded working set when the call arrives. When a serverless function restores from a snapshot, page faults dominate function time. REAP’s original number is 95%. Its method records the working set on the first call and prefetches that record on every later arrival. It depends on the working set being stable across calls. Seven of ten functions share more than 97% of their pages, the worst still shares 76%, and cold start drops by 3.7 times on average (Ustiugov et al. 2021, ASPLOS). FaaSnap points out that the working set moves sharply with the input, and REAP then gets worse. Concurrent faults let the guest start executing immediately, up to 3.5 times faster than REAP (Ao et al. 2022, EuroSys). Spice lays the snapshot out on disk in the predicted access order, so a restore from disk is only 0.6 to 18 milliseconds slower than a warm call (Holmes et al. 2026, OSDI). The usual base is a light microVM such as Firecracker, with under 5 MB of memory overhead per container (Agache et al. 2020, NSDI). This line makes restore fast, but the trigger is still call arrival, and the working sets are megabytes.

The third line starts before the call arrives. Serverless in the Wild records each function’s inter-arrival histogram and warms the function just before the next call (Shahrad et al. 2020, ATC). Orion uses the DAG structure plus a latency model to estimate when a downstream function will arrive, and warms the VM just before that (Mahgoub et al. 2022, OSDI). What they start early is a new instance. The signal is a statistic or a static structure, not an explicit announcement of one particular call before it happens. On the far-memory side, Leap prefetches by majority vote over trends in the access history (Al Maruf and Chowdhury 2020, ATC). That signal is history as well.

TL;DR

A tool call is visible in the inference stream before it has been fully generated. The natural guess is to restore the reclaimed sandbox during that lead, so wakeup leaves the critical path and idle sandboxes can sit near zero.

The lead is real. On 204 coding-agent sessions, and inside one KVM guest, we timed more than ten restore methods on a pandas sandbox, a Chromium sandbox, and a Vite sandbox, over zram and over a SATA disk. From the tool name until the call can be dispatched, the median is 1.71 seconds. It is not what dominates. On a 6.3 GB pandas sandbox on SATA, prefetching the previous call’s working set adds 6.1 seconds. Starting 1.7 seconds early only brings that to 4.7 seconds. Knowing exactly which pages this call will touch, even if the prefetch starts at arrival, brings it to 1.4 seconds, and the lead then brings it to 0.35 seconds.

Adjacent calls barely share pages. On pandas the overlap is 0.054. Restore cost follows how the pages sit on swap, not how many bytes they are. A call that touches the whole table takes 17 to 98 seconds under every method. In interactive sessions, 90.4% of idle memory-seconds sit in human waits longer than 600 seconds. Reclaim only there, and restore the whole sandbox when the request is issued. That already takes nine tenths of the memory benefit, at under 1.7% of active time. A new early signal has to be compared with the earliest signal the system already has. A restore stays hidden only when touched bytes fit inside the lead time times the effective bandwidth, and that bandwidth is set by the swap layout.

The signal that appears inside the inference stream

The serving side of LLMs is already using the fact that output is visible before it is finished. Anthropic’s streaming protocol emits the tool name when a tool-call block starts, and the arguments then stream out in pieces (Anthropic streaming documentation). LLMCompiler’s streaming planner starts executing tool calls that have already been parsed before the plan is finished (Kim et al. 2024, ICML). Pie splits the generation loop into fine-grained handlers given to the user program, so arbitrary computation and I/O can be inserted into the stream (Gim et al. 2025, SOSP). InferCept decides whether KV stays or goes while a tool call is in flight, but the restore happens after the tool returns (Abhyankar et al. 2024, ICML).

Put the two sides together and a position appears that nobody seems to occupy. A reclaimer has only been able to guess the next access from history. The inference stream announces that access several seconds before it happens. No layer hands that announcement to the memory manager.

An intuition: announce early, and the restore can hide

The intuition is a claim about the system. The start of reactivation can move from the moment of access to the moment the access is announced. The instant a tool name appears in the token stream, restore of that sandbox’s memory begins. By the time the arguments have streamed out and the call is actually dispatched, the pages are back in memory, and sandbox wakeup is no longer on the critical path. The reclaimer can then push the sandbox close to zero in every idle gap, without worrying about wakeup latency.

It looks plausible because three scales seem to line up. The lead time is seconds, because arguments stream out one token at a time. Reactivation is also seconds, because a working set of gigabytes comes back from disk or from compressed memory at a few gigabytes per second. A REAP-style prefetch that starts only when the call arrives depends on a stable working set. Each agent tool call has a different command and different files, so whether the working set is stable is exactly the unknown. An earlier trigger by itself is not enough. That is only a scheduling policy. For the claim to hold, two things have to be true at once. Reactivation is on the order of seconds and an arrival-time prefetch cannot absorb it, and the lead time is long enough and the working set can be predicted from the contents of the call.

Four falsification lines and the measurement setup

Four falsification lines were fixed before any measurement. The extra latency of an arrival-time prefetch has a median under 0.3 seconds, which would mean the strongest existing combination is already enough. Hiding all of that arrival-time extra latency saves less than 10% of per-step session time even on the most favorable sandbox. The lead-time median is under 0.5 seconds. The overlap of touch sets between adjacent calls is above 0.8, which would mean the working set is stable and REAP already covers it. If any one of these holds, the claim does not.

The lead time has two sources. The first is logs from 204 real coding-agent sessions from Claude Code, including sub-agents, with 13409 tool-call blocks. The log writes a line when each content block ends, so the start of a block is approximated by the timestamp of the previous content line in the same request, and the end of the block is the dispatch time. The lead time is the difference. This is a block-level approximation, not a token-level timestamp. The first block has no predecessor and is excluded, leaving 9599 blocks. The second source serves Qwen3.6-35B-A3B with vLLM 0.25.1 at tensor-parallel degree 2 on two A40s, and measures directly, at concurrency 1, 16, and 64, the time from when the tool name is visible on the streaming interface until the call completes.

Reactivation is measured inside one KVM guest: Linux 6.8, 16 vCPUs, 24 GiB of memory. Swap has two tiers. One is 20 GB of lz4 zram. The other is a 16 GB virtio disk on the host’s SATA SSD. After that disk is full, 4K random reads are about 23 MB/s at queue depth 1, about 354 MB/s at queue depth 32, and sequential reads are 536 MB/s. The three sandboxes are deliberately different materials. The pandas data-science sandbox is resident at 6.3 GB, almost all anonymous pages, and the data is a random floating-point dataframe. The Chromium browser-automation sandbox is 2.2 GB, 84% anonymous, spread across nine processes. The Vite plus TypeScript frontend-build sandbox is only 0.16 GB, mostly file pages. Each sandbox runs a real tool sequence of about ten calls. Before each call, memory.reclaim pushes the sandbox under 8 MiB. pagemap and mincore record the pages that call actually touches. Prefetch is issued by a helper process inside the sandbox’s control group, using process_madvise with MADV_WILLNEED and posix_fadvise (process_madvise manual).

The restore methods are laid out along signal and content. Demand paging. Prefetch of the previous call’s record at arrival, the REAP style. The same record started 1.7 seconds early. The exact touch set started at arrival. The exact touch set started 1.7 seconds early. And a full prefetch of the whole sandbox when the request is issued, 6.8 seconds early. The exact touch set is known only after the fact, so it is an upper bound on the dimension of knowing the content. Extra latency for every method is computed against the same call with the sandbox left resident and not reclaimed. Each method runs the call sequence once per sandbox. The resident baseline differs by up to a factor of two between the two passes, so differences under 0.5 seconds are inside the noise and cannot be ranked. As a comparison, CRIU does a process-level checkpoint, both a full restore and a lazy-pages restore. The latter is served on fault by a daemon through userfaultfd (CRIU lazy migration).

Evidence: the lead time is real, and it is not the dominant term

Timeline of one agent tool call. Restore can start when the request is issued, when the tool name is visible, or when the call is dispatched. Restore is fully hidden only if touched bytes do not exceed lead time times effective restore bandwidth.
Figure 1. One agent step, and the three moments a restore can start: when the request is issued, when the tool name becomes visible, and when the call is dispatched. A restore is fully hidden only if touched bytes stay within lead time times effective restore bandwidth.

Figure 1 is the reading path for the rest of the article. The top row is one agent step, from the request issued after a long gap, through the model’s reasoning text, through the tool name and arguments streaming out, to dispatch and execution. The three arrows are the three moments a restore can start: when the request is issued, which the agent’s execution framework already knows; when the tool name is visible, which is the claim above; and when the call is dispatched, which is demand paging and REAP. The box below is the condition every later number has to obey. Touched bytes must not exceed lead time times effective restore bandwidth, or the restore cannot be fully hidden. The two effective bandwidths in the box are read from the measurements. The pandas arrays are contiguous on swap: 546 MB comes back in 1.42 seconds, about 0.38 GB/s. Chromium’s heap pages are scattered: 148 MB takes 6.21 seconds, about 24 MB/s, close to the SATA disk’s 4K random-read speed at queue depth 1. The same 1.7 seconds of lead time hides about 0.65 GB in the first case and about 40 MB in the second.

CDF of lead time. In real sessions the median from tool-name visibility to dispatch is 1.71 seconds, and from request issue to dispatch is 6.84 seconds. On a local model the lead time rises from 0.50 to 2.82 seconds as concurrency grows.
Figure 2. Lead time before a call can be dispatched, on a log axis. From tool-name visibility the median is 1.71 seconds. From request issue it is 6.84 seconds. On a local model the lead time grows with load.

Figure 2’s horizontal axis is the lead time before a call can be dispatched, on a log scale, and the vertical axis is the cumulative fraction. The solid line is real sessions, from the tool name appearing until dispatch: median 1.71 seconds, p90 13.2 seconds, and only 4.2% below the 0.5-second falsification line, so that line does not fire. Its length is set by argument length, about 170 bytes per second. A token-level lead time is essentially the time it takes the arguments to stream out. The dashed line is from request issue until dispatch, median 6.84 seconds, four times the token-level lead. Of the requests issued immediately after a tool result, 93.0% end in a tool call. At the moment the request is issued, the execution framework already knows, with better than nine-tenths confidence, that the sandbox is about to be used. The three local-model lines show that lead time moves with load: 0.50 seconds at one stream, 1.09 seconds at 16, and 2.82 seconds at 64. The busier the inference, the longer the lead.

Extra latency of six restore methods against a resident baseline, median on the left and mean on the right, for three sandboxes and two swap tiers.
Figure 3. Extra latency against a resident sandbox. Knowing what to fetch matters more than knowing when. A call that touches the whole table is slow under every method.

Figure 3 is the decisive data. The horizontal axis is the six combinations of sandbox and tier. Each group has six bars, one per restore method. The vertical axis is extra latency against the resident baseline, on a log scale. Start with the most favorable group, the pandas sandbox on SATA. Demand paging adds a median of 4.04 seconds. Prefetching the set recorded from the previous call is worse: 6.11 seconds when started at arrival, and still 4.73 seconds when started 1.7 seconds early. The exact touch set, started at arrival, drops to 1.42 seconds, and adding the 1.7-second lead brings it to 0.35 seconds. Said another way, switching from history to exact content saves about 4.7 seconds. Adding the lead time on top of exact content saves only about one more second. That is where the title comes from. An arrival-time prefetch of 6.11 seconds is not under 0.3 seconds. Against a per-step average session time of 25.6 seconds, hiding it is 23.8%, above 10%. The first two falsification lines also do not fire.

Recorded prefetch is worse than demand paging because the touch sets of adjacent calls barely overlap. The median overlap is 0.054 on pandas, 0.22 to 0.26 on Chromium, and only the Vite build sandbox reaches 0.75. REAP measured 76% to 97% on serverless functions. In a data sandbox an agent reads different columns and different slices on each call. Most of the previous record is a wasted read, and that wasted read occupies disk bandwidth. The fourth falsification line, overlap above 0.8, also does not fire, and the data moves away from it.

The means on the right panel are a different fact. On pandas over disk, every method has a mean between 8 and 30 seconds, far above the median, pulled up by a few calls. Boolean filters, sorts, query, and sampling materialize the whole table, really touching 5.1 GB, and on SATA that takes 17 to 98 seconds no matter the method. Chromium on disk has a point that runs against intuition. Prefetching the exact 148 MB this call touches takes 6.2 seconds at arrival and still 5.4 seconds with a 1.7-second lead, while reading the whole 1.8 GB back sequentially when the request is issued adds only 0.45 seconds. Reading ten times the bytes is more than ten times faster. Restore cost is not paid per byte. It is paid by how scattered the swapped-out pages are. On zram every method stays within about one second, but random floating-point data compresses only to 1.23 times and saves only about 19% of memory. Chromium’s JavaScript heap compresses 3 to 13 times. That was measured on a synthetic page, and it is optimistic.

A few comparisons fill in the space outside the figure. A full CRIU restore of a 6.4 GB image has a median of 15.8 seconds. A lazy-pages restore itself takes only 0.55 seconds, but later calls slow down: eight calls average an extra 9.6 seconds, and one full-table sample under lazy-pages did not finish in more than 12 minutes, with the daemon pinned on one core. One reclaim plus a full swap-in drops transparent huge pages from 5990 MB to 22 MB, but the calls do not get slower. process_madvise with MADV_WILLNEED only reads pages into the swap cache. It does not build page tables. After a full prefetch, about 1.1 GB still takes a minor fault, a floor of 0.3 to 0.6 seconds.

Last is whether the content can be predicted before the call arrives. We wrote a rule that looks only at the call text. Column names that appear in the call imply which columns will be read, and those columns map to address ranges. The rule and 6 held-out calls were frozen before any held-out call ran. On 14 calls that do not materialize the whole table and whose touch set is nonempty, recall is 0.92 to 0.995 and precision is 1.0. Restoring from that prediction 1.7 seconds early has a median extra latency of 0.77 seconds on the original calls and 0.73 seconds on the held-out calls, against 4.2 to 4.3 seconds for demand paging. The 3 full-table calls have recall of only 0.09 to 0.19, so the mean returns to 25.6 seconds, almost level with demand paging at 29.2 seconds. If prefetch is issued column by column in the real streaming order of the arguments, the median is 1.1 to 1.4 seconds. Half of the timing is eaten by the arguments streaming out. The column name is itself inside the arguments. The later it appears, the less time is left for restore.

Tradeoff of four reclaim and restore policies on interactive sessions. Reclaiming only gaps longer than 600 seconds and restoring when the request is issued frees 90.4 percent of idle memory-seconds at under 1.7 percent of active time.
Figure 4. Four policies on real idle geometry. Reclaiming only gaps longer than 600 seconds, and restoring when the request is issued, frees 90.4% of idle memory-seconds. Delay stays under 1.7% of active time. Reclaiming every gap frees only 6 to 8 more points and costs one or two orders of magnitude more delay.

Figure 4 puts the measured restore cost back into the idle geometry of the real sessions. It is a desk calculation. Reclaim itself is charged at 2 seconds, which is optimistic for reclaim inside the tool loop. The horizontal axis is added delay as a fraction of the agent’s active time, on a log scale. The vertical axis is the fraction of sandbox idle memory-seconds that are freed. The circle reclaims only human-wait gaps longer than 600 seconds and restores the whole sandbox when the request is issued: 90.4% freed, 426 reactivations, and the delay on the worst case, pandas on disk, is 1.7% of active time, under 0.2% on the other sandboxes. The square is the policy that matches the claim above, reclaiming every gap and restoring with the 1.7-second token-level lead: 96.7% freed, 7442 reactivations, and on pandas over disk the delay is 20.0% with the exact set and 64.6% with the historical record. The triangle reclaims every gap and restores at arrival: 98.5% freed, with delay up to 111.7%. The 90.4% is the idle geometry itself. Across the 204 sessions, the short gaps inside the tool loop are only 9.6% of idle memory-seconds. Any reclaim scheme that acts only inside the tool loop has a memory-benefit ceiling of those 9.6 points.

Timing loses to content, and to an earlier signal

None of the four falsification lines that were fixed in advance fired. The claim is not stopped by one of them. The evidence forces a different statement of it, and that statement fails too.

Timing is not the dominant term. On the same sandbox and the same disk, switching from a historical record to exact content saves about 4.7 seconds. Adding a 1.7-second lead on top of exact content saves only about one more second. Without exact content, the lead time is almost useless. The claim treated timing as the valuable part. The measurements say the valuable part is which pages come back.

The token-level signal sits under a signal the system already has. When the request is issued, the execution framework already knows the sandbox is about to be used. That signal leads by a median of 6.84 seconds, four times the token-level lead, and 93% of those requests do end in a tool call. In interactive sessions, that request-level signal plus reclaim only during human-wait gaps already captures 90.4% of the memory benefit. The token-level scheme picks up only 6.3 more points and pays one or two orders of magnitude more delay.

Rewriting the claim as “predict the touch set from the call arguments” is still not enough. In an autonomous loop with no human-wait gap, for example an RL rollout, the request-level signal loses its advantage, and reclaim inside the loop is the only source of benefit. From the idle geometry, the ideal density ceiling of in-loop reclaim is about 1.76 times. That is the space this direction still has, and restoring from the arguments is the path that tries to take it. It holds only on calls that do not materialize the whole table. The full-table calls are exactly the ones that dominate the mean, and on SATA no lead time hides them. On zram the latency cost is only 1.8% to 3.8%, but numeric data saves only 19% of memory, a density of about 1.09 times. Wiring arguments to restore needs only a predictor. The mechanism itself can be assembled from a streaming interface, an Orion-style start before arrival, a REAP-style prefetch, and Linux’s reclaim interface. Take off the wrapper and what remains is a few new measurements, not a new mechanism. Work in the same direction, so far only on arXiv, warms a sandbox early from keywords and streaming embeddings in the token stream (SpecBox, arXiv 2607.23933). The trigger itself is not scarce.

A few gaps remain. The test machine has only a SATA disk. NVMe is about ten times faster, and a CXL pool is faster still, so the hiding boundary would move from 0.65 GB to several gigabytes and a full-table call might become hideable. That favors the claim, and it was not measured. All three sandboxes are constructed. The 6 GB pandas sandbox was chosen because it is the favorable case, and a real deployment’s sizes are unknown. Every sandbox number comes from a KVM guest, so faults include EPT overhead and restore looks more expensive than on bare metal. Each method ran the sequence once, so differences under 0.5 seconds cannot be ranked. Session lead times are timestamps of content blocks, not of tokens. Chromium’s compression ratio comes from a synthetic page. Sandboxes with a GPU were not measured.

Five measurements of agent sandbox memory

In real coding-agent sessions, the lead from when a tool name is visible until the call can be dispatched has a median of 1.71 seconds. Length follows the arguments, about 170 bytes per second. The lead from when the request is issued until dispatch has a median of 6.84 seconds, and 93% of those requests end in a tool call. A new early signal has to be compared with the earliest signal the system already has, not with having no signal at all. What it adds is precision and content. This was measured on 204 Claude Code sessions, at block-level approximation. The busier the inference, the longer the token-level lead. A local 35B mixture-of-experts model rises from 0.50 seconds at one stream to 2.82 seconds at 64.

The working set of an agent data sandbox does not come back on the next call. Overlap of touch sets between adjacent calls has a median of 0.054 on pandas and 0.22 to 0.26 on Chromium. Only the frontend-build sandbox reaches 0.75. Serverless functions sit at 76% to 97%. Recording the last call and prefetching it can therefore be worse than demand paging. On pandas over SATA the median extra latency is 6.1 seconds for the record and 4.0 seconds for demand paging. That is three constructed sandboxes, each with about ten calls. A sandbox dominated by file pages, such as a frontend build, is the exception.

Reactivation can be hidden when touched bytes do not exceed lead time times effective restore bandwidth. That bandwidth is set by the storage tier and by how the swapped-out pages are laid out, not by the device’s nominal bandwidth. On a SATA SSD, a contiguous array is about 0.38 GB/s. A browser heap scattered across nine processes is about 24 MB/s. The same 1.7 seconds hides 16 times more bytes in the first case. An exact prefetch of 148 MB on the Chromium sandbox takes 6.2 seconds, while a sequential read of the whole 1.8 GB adds only 0.45 seconds. A cost model that writes bytes divided by bandwidth underestimates by 3 to 16 times here.

In interactive agent sessions, 90.4% of sandbox idle memory-seconds sit in human-wait gaps longer than 600 seconds. The short gaps inside the tool loop are 9.6%. Reclaiming only the long gaps, and restoring the whole sandbox when the request is issued, captures nine tenths of the benefit. Even on the worst sandbox the delay is under 1.7% of active time. Any reclaim scheme that acts only inside the tool loop tops out at those 9.6 points. This depends on how often a person comes back in these sessions, and on charging reclaim at 2 seconds. It does not hold in an autonomous loop with no person in it. There, reclaim inside the loop is the only source of benefit.

At column granularity, the pages one call touches in a data sandbox can be predicted from the call text, with recall above 0.92 and precision 1.0. That holds only for calls that do not materialize the whole table. Boolean filters, sorts, query, and sampling touch the whole table, and recall falls below 0.2. On SATA every restore method then takes 17 to 98 seconds. The calls that can be predicted were cheap anyway. The expensive ones cannot be predicted. This was measured on a 6 GB random floating-point dataframe and on frozen held-out calls. A faster storage tier would shrink the cost of the full-table calls.

The coverage of this article stops here: one KVM guest, SATA and zram, three constructed sandboxes, and one batch of coding-agent sessions. It does not show that a token-level lead time can move sandbox reactivation off the critical path. It does not show that agent sandbox working sets are generally gigabytes. It does not extend to NVMe or CXL. What it does show, under these conditions, is that the cost of waking a sandbox is what has to come back and how it is laid out, not when the fetch starts.

Ninety percent of idle memory is the human wait

In interactive sessions, 90.4% of sandbox idle memory-seconds sit in human-wait gaps longer than 600 seconds. Reclaim only there, and restore the whole sandbox when the request is issued. That already takes nine tenths of the memory benefit, at a delay under 1.7% of active time. A new early signal has to be compared with the earliest signal the system already has. Whether a restore can be hidden depends on how many bytes are touched, how long the lead is, and the effective restore bandwidth. That bandwidth is set by how the pages sit on swap.

Agent Sandbox: What Can Still Cut Runtime Memory and Page Faults looks at the same wakeup from the other side. After reclaim, where the pages sit on swap, and how much of that disk read a placement change can still remove on Linux 6.1 and 6.12.