How Much Memory Can an Agent Sandbox Swarm Still Share?

Once the repo is shared read-only, a same-repo sandbox swarm has only 4% to 10% of its memory left to deduplicate, and almost all of that remainder is private heaps, not files.

A coding agent used to be one session in one sandbox. It is becoming a swarm. Subagents hand subtasks to parallel environments. Best-of-N tries N edits of the same problem at once. Reinforcement-learning rollouts run hundreds or thousands of trajectories on one repository snapshot. They share one fact: on the same host, these environments read the same repository, the same dependencies, and the same compiler and language runtime. The natural question is whether the host is holding many copies of identical pages, and whether that duplication is what caps how many sandboxes one machine can run. This post is how we measured that, and why the measurement leaves a new shared-memory abstraction with almost nothing to name.

Where a page of memory comes from

Inside sibling sandboxes on the same repository, file pages and anonymous pages are shared by completely different mechanisms.

flowchart LR
  Repo["shared repo"] --> PC["host page cache<br/>one inode, one copy"]
  Tool["tsserver / pytest"] --> A1["heap A"]
  Tool --> A2["heap B"]
  Tool --> A3["heap C"]
  PC --> S["file pages: usually one copy"]
  A1 --> D["duplicates live in anon heaps"]
  A2 --> D
  A3 --> D

The question is not whether sandboxes repeat memory in general. After the file layer is already shared, how many identical anonymous pages are left that would justify a separate cross-sandbox dedup?

Linux physical memory is counted in 4 KiB pages, and a process sees two kinds. A file page is a cache of file contents in the kernel page cache. Reads, executable mappings, and shared libraries land there. One inode and one offset exist as one physical page on the host, no matter how many processes are reading. An anonymous page has no file behind it. The heap, the stack, JIT code, and language-runtime objects live there, and each process gets its own by default.

Containers share file pages through overlayfs. It stacks a read-only lower directory and a writable upper directory into one view. A read of an unmodified file hits the same inode in the lower layer. A write copies the whole file up into the private upper layer (Linux overlayfs). A hundred containers from one image can read one library file and still occupy one page-cache copy. The cost is write amplification by file, not by page. Changing one byte of a large file copies the whole file.

A virtual machine does not get that. A microVM guest has its own kernel and its own page cache. When it reads a block image through virtio-blk, the bytes enter the host page cache first, unless the host used O_DIRECT, and are then copied into the guest page cache. From the host, that second copy is anonymous memory of the hypervisor process. N guests reading one image means N guest page caches. Firecracker is a minimal VMM for serverless, with a device model limited to virtio-net, virtio-blk, vsock, and a few others (Agache et al., 2020). The way to remove the guest-side copy is DAX, direct access. A virtiofs DAX window, or a virtio-pmem device, maps the host file straight into the guest physical address space. A guest filesystem mounted with DAX builds no page cache of its own and reads the host’s copy (virtio-fs). Firecracker has supported virtio-pmem since 1.14. The backing file is mapped MAP_SHARED, can be read-only, and a DAX mount in the guest bypasses the guest page cache. The same note warns that several VMs sharing one backing file are a cross-VM side channel (Firecracker pmem).

Two other kinds of sharing happen at different times. One is share-at-birth. fork lets the child share every page with the parent until a write. A zygote or a fork server warms the runtime and then forks it out to each task. Firecracker snapshot restore maps the memory file MAP_PRIVATE, so several VMs restored from one snapshot share the clean pages until they diverge. REAP records the working set actually touched after restore and faults it in as a batch (Ustiugov et al., 2021). FaaSnap lays the snapshot out in page-access order (Ao et al., 2022). Catalyzer derives a new instance from an initialized template sandbox with sfork (Du et al., 2020). They mainly cut cold-start time, and memory sharing at birth comes along. The other kind is merge-after-the-fact. KSM periodically scans anonymous regions marked MADV_MERGEABLE and collapses identical pages into one copy-on-write page, at the cost of scan CPU, the peak before the merge, and a known side channel (Linux KSM). Medes takes that idea into a serverless platform and deduplicates warm-instance memory across sandboxes by content (Saxena et al., 2022).

Together these cover three intervals: share at birth (fork, snapshots, zygote), share on read (overlayfs, DAX), and merge later (KSM, Medes). Every fraction below has to be read against that combination, not against a machine that shares nothing.

A resident tsserver on a medium TypeScript repository is about 0.5 GB, and more than nine tenths of that is anonymous. A minimal Node 20 plus Python toolchain image is about 0.2 GB of files, and less than 0.15 GB of that is actually read into the page cache. In a sibling’s resident set, the file part was small to begin with.

Why a large duplicate looked plausible

The abstraction we started from is concrete. The host keeps a content-addressed namespace of read-only page objects. Each version of the repo, the dependencies, and the toolchain is a set of immutable objects. Every sibling maps the same set from birth. A write affects only the writer. Switching versions is switching object sets. The intuition is that sibling read sets overlap heavily, so the waste of not sharing grows linearly with the swarm.

Checking that idea against existing mechanisms, before building anything, already showed the risk. On the file-page side, overlayfs, DAX, and KSM together already express share-at-birth and privatize-on-write. What remains is pages that happen to match inside mutable anonymous heaps, and those are not immutable objects. A memory census of developer-tool processes from an earlier round found that anonymous memory is 91% to 96% of the resident set of tsserver, a resident pytest worker, and a javac daemon. At that ratio, page-level dedup of 32 siblings saves another 6.6% to 9.4% beyond a DAX configuration, under the 20% gain we had set as the bar.

That was an estimate. It assumes the anonymous fraction inside a swarm matches a single process, and that the guest holds no large read-only derived data that DAX cannot cover. Both could be wrong. When a swarm reads one repository together, read buffers, parsed source strings, and build products might each sit as identical bytes in a private heap. That is immutable content that has entered anonymous memory, which file sharing cannot see, and it is exactly where a new abstraction might still pay. So we measured it.

TL;DR

Subagents, best-of-N, and RL rollouts put several to several dozen sandboxes on one repository. They read one repo, one dependency set, and one toolchain, so it is natural to expect the host to be storing identical pages more than once, and to want a content-addressed read-only namespace shared from birth.

On one 112-core server we ran 4 to 16 siblings for 10 minutes against one TypeScript repository, with a resident tsserver, grep, tsc, pytest, and a private file edit each. A page-content census compared four setups with a swarm-wide dedup oracle: host processes, containers sharing a read-only repo, containers that each copy the repo, and Firecracker with no sharing.

Once the repo and toolchain come in as a read-only layer, the unshared duplicate after 10 minutes is 4% to 10% of the swarm’s memory. About 96% of that duplicate is each sibling’s private anonymous heap. The file pages are already one copy. Of the heap duplicate, 52% to 92% already exists at birth, from the same tools starting the same way. Only Firecracker with no sharing leaves 37% to 42% duplicate, mostly each guest’s own page cache and kernel pages, which is what virtio-pmem plus DAX is built to remove. In the strong configuration, that namespace has almost nothing left to name.

How the census was done

Shared read-only layer versus Firecracker with no sharing. The first leaves private heaps. The second gives every guest its own page cache and kernel.
Figure 1. With a shared read-only layer, each sibling keeps only a private heap. A microVM with no sharing gives every guest its own page cache and kernel.

Figure 1 is the two extremes. On the left, containers or host processes read one read-only layer. The repo and toolchain exist once on the host, and what each sibling owns is the heap of its own tool processes. On the right, Firecracker mounts one rootfs image read-only through virtio-blk. The image exists once in the host page cache, but each guest’s memory also holds its own page-cache copy and its own kernel.

The machine is one x86 server with 112 cores and 251 GB of memory, kernel 5.15, using only the CPU and KVM. The workload is a fixed commit of date-fns, a TypeScript repository built with TypeScript 5.9, plus a small Python library, more-itertools. Each sibling runs for 10 minutes. A resident tsserver opens 200 source files and serves diagnostics, quickinfo, references, and completions. Every round greps the repository, edits one file owned by that sibling, and pushes the edit to tsserver. Every third round runs a full tsc typecheck. Every second round runs pytest. File choice is seeded by the sibling’s index, so the work is similar and not identical.

There are four configurations. In the host-process arm, each sibling is a group of host processes reading one work directory through its own overlayfs. The shared read-only container arm imitates Docker: separate mount, pid, ipc, and uts namespaces, chrooted into a root stacked from one image lower layer, with the repo mounted as a read-only lower plus a private upper. The copied-repo container arm is the same, except each sibling has its own copy of the repo. The Firecracker arm uses 1.10.1. Each VM has 2 vCPUs and 1.5 GiB, mounts one ext4 image read-only, and runs the same sibling script in the guest. This Firecracker build has no virtio-pmem, so the DAX arm was not measured. That is the most important gap in the post.

The census runs at 15 seconds, 1 minute, 5 minutes, and 10 minutes. Each sample freezes every sibling with a cgroup freeze, reads each process pagemap, takes the resident anonymous pages, deduplicates them by physical page number to get the real footprint, and hashes every page. For every file the swarm can see, mincore finds the pages in the page cache. Real occupancy counts those pages by inode. The oracle counts them by content hash. Real occupancy is file pages plus deduplicated anonymous pages. The oracle is the number of distinct contents across the swarm’s file and anonymous pages. The difference is content that matches and is not shared. That duplicate splits three ways: file pages with the same content in different inodes; anonymous pages whose content equals some immutable file block, which covers the guest page cache, read buffers, and file bytes copied into a heap; and the rest of the anonymous duplicate, pages that merely happen to match inside private heaps. The third class is split again by whether the content already appears in the 15-second snapshot. In the Firecracker arm, all guest memory is anonymous memory of the VMM process, so the guest page cache falls in the second class. Each cell is one run.

N is 4, 8, and 16. The shared read-only container arm and the Firecracker arm did not finish N=16, and no arm finished N=32. At N=4 and N=8 those two arms were already far from the threshold, and we stopped the probe there.

What the census found

Unshared duplicate memory by configuration. Shared read-only setups leave 6 to 8 percent, almost all private heap. Copied repos and unshared VMs duplicate files and guest page cache.
Figure 2. The two shared read-only setups leave 6% to 8% duplicate, almost all of it private heap. Copying the repo, or running an unshared VM, duplicates files and the guest page cache.

The left of Figure 2 is unshared duplicate memory at 10 minutes, for each configuration and each N. Bar height is gigabytes. The label is that duplicate as a fraction of the swarm’s actual memory. Color is the source: file page cache, anonymous pages equal to an immutable file block, private heap already present at the start, and private heap that diverged later. The right is, at the largest N of each configuration, the average actual footprint per sibling in light gray and the oracle in dark gray.

On the host-process arm and the shared read-only container arm, across every one of the 15 samples from N=4 to 16 and from 1 to 10 minutes, the duplicate stays between 3.8% and 9.8%. At 10 minutes it is 5.6% to 8.2%. The color is almost all gray and green, which is the private heap: 96% to 97% on the host arm, 90% to 92% on the container arm. File-page duplicate is 0% to 2%. The repo and toolchain each sibling reads really do exist once on the host. Total file pages stay near 0.15 GB on the host arm and 0.11 GB on the container arm from N=4 to 16, and do not grow with N. What grows linearly with N is anonymous memory. Of 6.76 GB on the host arm at N=16, 6.61 GB is anonymous.

The class we worried about most, immutable file content copied into each heap, is 3% to 9% of the duplicate in these two configurations, a few thousandths of swarm memory. After tsserver and node read the source, what they keep is an AST, a symbol table, and type information, not the original 4 KiB pages.

Copying the repo pushes the duplicate to 16% to 21%, and about 70% of that is file pages. Each sibling has its own repo inodes, so identical sources sit once per copy in the page cache. This is the configuration an overlayfs lower layer exists to avoid. The problem is real. It already has a standard fix.

Firecracker with no sharing has the largest duplicate: 38.5% to 38.8% at 10 minutes, and up to 41.5% at 1 minute. Of that, 17% to 29% is file pages, the part of the host page cache whose content matches the guest cache. Another 39% to 47% is guest memory whose content equals an image block, which is the guest page cache. The rest is matching pages in the guest’s private heap and kernel. On the right, each VM averages 0.62 GB and the oracle needs 0.38 GB. This is the only configuration over the 20% bar, and the excess is mostly the guest page cache. virtio-pmem plus DAX maps the guest onto the host’s copy and removes exactly that piece. We did not measure the DAX arm, so we cannot say what remains after it. From the makeup of the shared read-only container arm, what would remain is mostly heaps and the guest kernel.

The heap duplicates are born at startup

The hatched gray part of each bar in Figure 2 is private-heap duplicate that already exists in the 15-second snapshot. Across all 24 samples after 1 minute, it is 52% to 92% of the private-heap duplicate, and 61% to 92% on the host and shared read-only container arms. The heap pages that match across siblings are mostly not something the run grew into by accident. They are what the same node and tsserver binary, and the same Python interpreter, produce after startup and initialization on the same inputs. Counting how many siblings hold each matching page, the three non-VM configurations have about 3200 to 3300 pages, about 13 MB, held by every sibling at every N. The count barely moves with N. It is the part every tool startup generates.

That piece already has an owner. A zygote, a fork server, and snapshot restore make those pages one copy at birth, and a later write privatizes them. In the earlier census of the same class of tools, replaying one request sequence left only 3% to 12% of pages identical to the original process. GC, allocation order, and hash seeds pull two heaps apart in bytes even when the answers match. In this swarm, matching pages that appear only after startup are under 2% of swarm memory, which lines up.

Why the conjecture does not hold

The hypothesis needs two things at once. Under a strong sharing configuration, a large set of identical pages is still unshared, and those pages are immutable content, fit to be objects shared from birth. The first fails in the strong configuration. The remainder is 4% to 10%, less than half of the 20% bar. The second fails too. More than nine tenths of the remainder is a private mutable heap, which by definition sits outside an immutable object namespace. The part that exists at birth is already taken by fork and snapshots. The only configuration with a large opening is the microVM swarm that shares nothing, and that opening is the guest page cache. DAX and virtio-pmem are already the standard fix. The decisive row is the shared read-only container arm. That is Docker’s default, and it needs no new mechanism.

A few gaps remain. Each cell ran once. The shared read-only container arm and the Firecracker arm have no N=16 or N=32. On the host arm the duplicate fraction does not rise with N from 4 to 16. It is 5.6%, 8.2%, and 5.9%. We extrapolate that N=32 would not cross the bar, and that is an extrapolation. The workload is mostly read-only analysis of a medium repository. A multi-gigabyte node_modules or a large monorepo would enlarge the file pages, but the part that grows is the part already shared, which favors the conclusion. Sessions longer than 10 minutes, several repositories on one host, and large read-only derived products sitting in private heaps are not covered. Indexes, build caches, and vector stores are the last of those. They could raise the fraction of immutable content that has entered anonymous memory, and that is where this conclusion is most likely to fail.

Where the sharing opportunity actually remains

In a same-repo sandbox swarm, once the repo and toolchain are a shared read-only layer (a host overlayfs, or a container image lower layer), the unshared duplicate at 10 minutes is 4% to 10% of actual memory. About nine tenths of it is each sibling’s private anonymous heap. That is 4 to 16 siblings, a medium TypeScript and Python repository, a resident language server plus tests and builds, 4 KiB pages, and one run per cell.

What grows linearly with the number of siblings is anonymous memory. The file-page total barely moves. A developer tool’s resident set is already mostly heap, so the memory of a swarm is each sibling’s tool-process heap, not a file stored more than once. The conditions are the same as above.

Of the private-heap pages that match across siblings, 52% to 92% are already there at birth, from the same tools starting the same way. Once the read-only layer is shared, matching pages that appear only later are under 2% of swarm memory. Saving that part means sharing at birth, with a zygote, a fork server, or a snapshot restore, not content addressing or a scan during the run. The tool binary and the inputs have to match across siblings.

A microVM swarm that shares nothing leaves 37% to 42% duplicate. Most of it is each guest’s own page cache and identical kernel pages. virtiofs DAX and virtio-pmem exist to remove the guest page cache. A claim about how much a microVM swarm can still share has to include DAX in the baseline, or it will book DAX’s gain as its own. What remains after DAX was not measured.

The most a sharing mechanism can save is the slice of the remaining duplicate, after the strongest sharing that already exists, that belongs to the class of object it names. This census splits that remainder into four classes: files, immutable content that entered a heap, heap pages present at birth, and heap pages grown during the run. Any design that claims to deduplicate across sandboxes can first ask which class it owns, and how large that class is.

Subtract the file sharing that already exists

A same-repo sandbox swarm looks as if it must copy a lot. The Linux page cache, overlayfs, and a shared read-only layer have already taken the part that is easy to share. What remains is mostly private anonymous heap, and under a strong configuration it is a small slice of total memory. The first step in judging whether a new memory-sharing mechanism has an object is to subtract the file sharing that already exists, then ask how much of the anonymous remainder is duplicate, safe to share, and still there later.