How Much Memory Can an Agent Sandbox Swarm Still Share?
Once the repo is shared read-only, a same-repo sandbox swarm has only 4% to 10% of its memory left to deduplicate, and almost all of that remainder is private heaps, not files.
A coding agent used to be one session in one sandbox. It is becoming a swarm. Subagents hand subtasks to parallel environments. Best-of-N tries N edits of the same problem at once. Reinforcement-learning rollouts run hundreds or thousands of trajectories on one repository snapshot. They share one fact: on the same host, these environments read the same repository, the same dependencies, and the same compiler and language runtime. The natural question is whether the host is holding many copies of identical pages, and whether that duplication is what caps how many sandboxes one machine can run. This post is how we measured that, and why the measurement leaves a new shared-memory abstraction with almost nothing to name.
Where a page of memory comes from
Inside sibling sandboxes on the same repository, file pages and anonymous pages are shared by completely different mechanisms.
flowchart LR
Repo["shared repo"] --> PC["host page cache<br/>one inode, one copy"]
Tool["tsserver / pytest"] --> A1["heap A"]
Tool --> A2["heap B"]
Tool --> A3["heap C"]
PC --> S["file pages: usually one copy"]
A1 --> D["duplicates live in anon heaps"]
A2 --> D
A3 --> D
The question is not whether sandboxes repeat memory in general. After the file layer is already shared, how many identical anonymous pages are left that would justify a separate cross-sandbox dedup?
Linux physical memory is counted in 4 KiB pages, and a process sees two kinds. A file page is a cache of file contents in the kernel page cache. Reads, executable mappings, and shared libraries land there. One inode and one offset exist as one physical page on the host, no matter how many processes are reading. An anonymous page has no file behind it. The heap, the stack, JIT code, and language-runtime objects live there, and each process gets its own by default.
Containers share file pages through overlayfs. It stacks a read-only lower directory and a writable upper directory into one view. A read of an unmodified file hits the same inode in the lower layer. A write copies the whole file up into the private upper layer (Linux overlayfs). A hundred containers from one image can read one library file and still occupy one page-cache copy. The cost is write amplification by file, not by page. Changing one byte of a large file copies the whole file.
A virtual machine does not get that. A microVM guest has its own kernel and its own page cache. When it reads a block image through virtio-blk, the bytes enter the host page cache first, unless the host used O_DIRECT, and are then copied into the guest page cache. From the host, that second copy is anonymous memory of the hypervisor process. N guests reading one image means N guest page caches. Firecracker is a minimal VMM for serverless, with a device model limited to virtio-net, virtio-blk, vsock, and a few others (Agache et al., 2020). The way to remove the guest-side copy is DAX, direct access. A virtiofs DAX window, or a virtio-pmem device, maps the host file straight into the guest physical address space. A guest filesystem mounted with DAX builds no page cache of its own and reads the host’s copy (virtio-fs). Firecracker has supported virtio-pmem since 1.14. The backing file is mapped MAP_SHARED, can be read-only, and a DAX mount in the guest bypasses the guest page cache. The same note warns that several VMs sharing one backing file are a cross-VM side channel (Firecracker pmem).
Two other kinds of sharing happen at different times. One is share-at-birth. fork lets the child share every page with the parent until a write. A zygote or a fork server warms the runtime and then forks it out to each task. Firecracker snapshot restore maps the memory file MAP_PRIVATE, so several VMs restored from one snapshot share the clean pages until they diverge. REAP records the working set actually touched after restore and faults it in as a batch (Ustiugov et al., 2021). FaaSnap lays the snapshot out in page-access order (Ao et al., 2022). Catalyzer derives a new instance from an initialized template sandbox with sfork (Du et al., 2020). They mainly cut cold-start time, and memory sharing at birth comes along. The other kind is merge-after-the-fact. KSM periodically scans anonymous regions marked MADV_MERGEABLE and collapses identical pages into one copy-on-write page, at the cost of scan CPU, the peak before the merge, and a known side channel (Linux KSM). Medes takes that idea into a serverless platform and deduplicates warm-instance memory across sandboxes by content (Saxena et al., 2022).
Together these cover three intervals: share at birth (fork, snapshots, zygote), share on read (overlayfs, DAX), and merge later (KSM, Medes). Every fraction below has to be read against that combination, not against a machine that shares nothing.
A resident tsserver on a medium TypeScript repository is about 0.5 GB, and more than nine tenths of that is anonymous. A minimal Node 20 plus Python toolchain image is about 0.2 GB of files, and less than 0.15 GB of that is actually read into the page cache. In a sibling’s resident set, the file part was small to begin with.
Why a large duplicate looked plausible
The abstraction we started from is concrete. The host keeps a content-addressed namespace of read-only page objects. Each version of the repo, the dependencies, and the toolchain is a set of immutable objects. Every sibling maps the same set from birth. A write affects only the writer. Switching versions is switching object sets. The intuition is that sibling read sets overlap heavily, so the waste of not sharing grows linearly with the swarm.
Checking that idea against existing mechanisms, before building anything, already showed the risk. On the file-page side, overlayfs, DAX, and KSM together already express share-at-birth and privatize-on-write. What remains is pages that happen to match inside mutable anonymous heaps, and those are not immutable objects. A memory census of developer-tool processes from an earlier round found that anonymous memory is 91% to 96% of the resident set of tsserver, a resident pytest worker, and a javac daemon. At that ratio, page-level dedup of 32 siblings saves another 6.6% to 9.4% beyond a DAX configuration, under the 20% gain we had set as the bar.
That was an estimate. It assumes the anonymous fraction inside a swarm matches a single process, and that the guest holds no large read-only derived data that DAX cannot cover. Both could be wrong. When a swarm reads one repository together, read buffers, parsed source strings, and build products might each sit as identical bytes in a private heap. That is immutable content that has entered anonymous memory, which file sharing cannot see, and it is exactly where a new abstraction might still pay. So we measured it.
TL;DR
Subagents, best-of-N, and RL rollouts put several to several dozen sandboxes on one repository. They read one repo, one dependency set, and one toolchain, so it is natural to expect the host to be storing identical pages more than once, and to want a content-addressed read-only namespace shared from birth.
On one 112-core server we ran 4 to 16 siblings for 10 minutes against one TypeScript repository, with a resident tsserver, grep, tsc, pytest, and a private file edit each. A page-content census compared four setups with a swarm-wide dedup oracle: host processes, containers sharing a read-only repo, containers that each copy the repo, and Firecracker with no sharing.
Once the repo and toolchain come in as a read-only layer, the unshared duplicate after 10 minutes is 4% to 10% of the swarm’s memory. About 96% of that duplicate is each sibling’s private anonymous heap. The file pages are already one copy. Of the heap duplicate, 52% to 92% already exists at birth, from the same tools starting the same way. Only Firecracker with no sharing leaves 37% to 42% duplicate, mostly each guest’s own page cache and kernel pages, which is what virtio-pmem plus DAX is built to remove. In the strong configuration, that namespace has almost nothing left to name.
How the census was done
Figure 1 is the two extremes. On the left, containers or host processes read one read-only layer. The repo and toolchain exist once on the host, and what each sibling owns is the heap of its own tool processes. On the right, Firecracker mounts one rootfs image read-only through virtio-blk. The image exists once in the host page cache, but each guest’s memory also holds its own page-cache copy and its own kernel.
The machine is one x86 server with 112 cores and 251 GB of memory, kernel 5.15, using only the CPU and KVM. The workload is a fixed commit of date-fns, a TypeScript repository built with TypeScript 5.9, plus a small Python library, more-itertools. Each sibling runs for 10 minutes. A resident tsserver opens 200 source files and serves diagnostics, quickinfo, references, and completions. Every round greps the repository, edits one file owned by that sibling, and pushes the edit to tsserver. Every third round runs a full tsc typecheck. Every second round runs pytest. File choice is seeded by the sibling’s index, so the work is similar and not identical.
There are four configurations. In the host-process arm, each sibling is a group of host processes reading one work directory through its own overlayfs. The shared read-only container arm imitates Docker: separate mount, pid, ipc, and uts namespaces, chrooted into a root stacked from one image lower layer, with the repo mounted as a read-only lower plus a private upper. The copied-repo container arm is the same, except each sibling has its own copy of the repo. The Firecracker arm uses 1.10.1. Each VM has 2 vCPUs and 1.5 GiB, mounts one ext4 image read-only, and runs the same sibling script in the guest. This Firecracker build has no virtio-pmem, so the DAX arm was not measured. That is the most important gap in the post.
The census runs at 15 seconds, 1 minute, 5 minutes, and 10 minutes. Each sample freezes every sibling with a cgroup freeze, reads each process pagemap, takes the resident anonymous pages, deduplicates them by physical page number to get the real footprint, and hashes every page. For every file the swarm can see, mincore finds the pages in the page cache. Real occupancy counts those pages by inode. The oracle counts them by content hash. Real occupancy is file pages plus deduplicated anonymous pages. The oracle is the number of distinct contents across the swarm’s file and anonymous pages. The difference is content that matches and is not shared. That duplicate splits three ways: file pages with the same content in different inodes; anonymous pages whose content equals some immutable file block, which covers the guest page cache, read buffers, and file bytes copied into a heap; and the rest of the anonymous duplicate, pages that merely happen to match inside private heaps. The third class is split again by whether the content already appears in the 15-second snapshot. In the Firecracker arm, all guest memory is anonymous memory of the VMM process, so the guest page cache falls in the second class. Each cell is one run.
N is 4, 8, and 16. The shared read-only container arm and the Firecracker arm did not finish N=16, and no arm finished N=32. At N=4 and N=8 those two arms were already far from the threshold, and we stopped the probe there.
What the census found
The left of Figure 2 is unshared duplicate memory at 10 minutes, for each configuration and each N. Bar height is gigabytes. The label is that duplicate as a fraction of the swarm’s actual memory. Color is the source: file page cache, anonymous pages equal to an immutable file block, private heap already present at the start, and private heap that diverged later. The right is, at the largest N of each configuration, the average actual footprint per sibling in light gray and the oracle in dark gray.
On the host-process arm and the shared read-only container arm, across every one of the 15 samples from N=4 to 16 and from 1 to 10 minutes, the duplicate stays between 3.8% and 9.8%. At 10 minutes it is 5.6% to 8.2%. The color is almost all gray and green, which is the private heap: 96% to 97% on the host arm, 90% to 92% on the container arm. File-page duplicate is 0% to 2%. The repo and toolchain each sibling reads really do exist once on the host. Total file pages stay near 0.15 GB on the host arm and 0.11 GB on the container arm from N=4 to 16, and do not grow with N. What grows linearly with N is anonymous memory. Of 6.76 GB on the host arm at N=16, 6.61 GB is anonymous.
The class we worried about most, immutable file content copied into each heap, is 3% to 9% of the duplicate in these two configurations, a few thousandths of swarm memory. After tsserver and node read the source, what they keep is an AST, a symbol table, and type information, not the original 4 KiB pages.
Copying the repo pushes the duplicate to 16% to 21%, and about 70% of that is file pages. Each sibling has its own repo inodes, so identical sources sit once per copy in the page cache. This is the configuration an overlayfs lower layer exists to avoid. The problem is real. It already has a standard fix.
Firecracker with no sharing has the largest duplicate: 38.5% to 38.8% at 10 minutes, and up to 41.5% at 1 minute. Of that, 17% to 29% is file pages, the part of the host page cache whose content matches the guest cache. Another 39% to 47% is guest memory whose content equals an image block, which is the guest page cache. The rest is matching pages in the guest’s private heap and kernel. On the right, each VM averages 0.62 GB and the oracle needs 0.38 GB. This is the only configuration over the 20% bar, and the excess is mostly the guest page cache. virtio-pmem plus DAX maps the guest onto the host’s copy and removes exactly that piece. We did not measure the DAX arm, so we cannot say what remains after it. From the makeup of the shared read-only container arm, what would remain is mostly heaps and the guest kernel.
The heap duplicates are born at startup
The hatched gray part of each bar in Figure 2 is private-heap duplicate that already exists in the 15-second snapshot. Across all 24 samples after 1 minute, it is 52% to 92% of the private-heap duplicate, and 61% to 92% on the host and shared read-only container arms. The heap pages that match across siblings are mostly not something the run grew into by accident. They are what the same node and tsserver binary, and the same Python interpreter, produce after startup and initialization on the same inputs. Counting how many siblings hold each matching page, the three non-VM configurations have about 3200 to 3300 pages, about 13 MB, held by every sibling at every N. The count barely moves with N. It is the part every tool startup generates.
That piece already has an owner. A zygote, a fork server, and snapshot restore make those pages one copy at birth, and a later write privatizes them. In the earlier census of the same class of tools, replaying one request sequence left only 3% to 12% of pages identical to the original process. GC, allocation order, and hash seeds pull two heaps apart in bytes even when the answers match. In this swarm, matching pages that appear only after startup are under 2% of swarm memory, which lines up.
Why the conjecture does not hold
The hypothesis needs two things at once. Under a strong sharing configuration, a large set of identical pages is still unshared, and those pages are immutable content, fit to be objects shared from birth. The first fails in the strong configuration. The remainder is 4% to 10%, less than half of the 20% bar. The second fails too. More than nine tenths of the remainder is a private mutable heap, which by definition sits outside an immutable object namespace. The part that exists at birth is already taken by fork and snapshots. The only configuration with a large opening is the microVM swarm that shares nothing, and that opening is the guest page cache. DAX and virtio-pmem are already the standard fix. The decisive row is the shared read-only container arm. That is Docker’s default, and it needs no new mechanism.
A few gaps remain. Each cell ran once. The shared read-only container arm and the Firecracker arm have no N=16 or N=32. On the host arm the duplicate fraction does not rise with N from 4 to 16. It is 5.6%, 8.2%, and 5.9%. We extrapolate that N=32 would not cross the bar, and that is an extrapolation. The workload is mostly read-only analysis of a medium repository. A multi-gigabyte node_modules or a large monorepo would enlarge the file pages, but the part that grows is the part already shared, which favors the conclusion. Sessions longer than 10 minutes, several repositories on one host, and large read-only derived products sitting in private heaps are not covered. Indexes, build caches, and vector stores are the last of those. They could raise the fraction of immutable content that has entered anonymous memory, and that is where this conclusion is most likely to fail.
Where the sharing opportunity actually remains
In a same-repo sandbox swarm, once the repo and toolchain are a shared read-only layer (a host overlayfs, or a container image lower layer), the unshared duplicate at 10 minutes is 4% to 10% of actual memory. About nine tenths of it is each sibling’s private anonymous heap. That is 4 to 16 siblings, a medium TypeScript and Python repository, a resident language server plus tests and builds, 4 KiB pages, and one run per cell.
What grows linearly with the number of siblings is anonymous memory. The file-page total barely moves. A developer tool’s resident set is already mostly heap, so the memory of a swarm is each sibling’s tool-process heap, not a file stored more than once. The conditions are the same as above.
Of the private-heap pages that match across siblings, 52% to 92% are already there at birth, from the same tools starting the same way. Once the read-only layer is shared, matching pages that appear only later are under 2% of swarm memory. Saving that part means sharing at birth, with a zygote, a fork server, or a snapshot restore, not content addressing or a scan during the run. The tool binary and the inputs have to match across siblings.
A microVM swarm that shares nothing leaves 37% to 42% duplicate. Most of it is each guest’s own page cache and identical kernel pages. virtiofs DAX and virtio-pmem exist to remove the guest page cache. A claim about how much a microVM swarm can still share has to include DAX in the baseline, or it will book DAX’s gain as its own. What remains after DAX was not measured.
The most a sharing mechanism can save is the slice of the remaining duplicate, after the strongest sharing that already exists, that belongs to the class of object it names. This census splits that remainder into four classes: files, immutable content that entered a heap, heap pages present at birth, and heap pages grown during the run. Any design that claims to deduplicate across sandboxes can first ask which class it owns, and how large that class is.
Subtract the file sharing that already exists
A same-repo sandbox swarm looks as if it must copy a lot. The Linux page cache, overlayfs, and a shared read-only layer have already taken the part that is easy to share. What remains is mostly private anonymous heap, and under a strong configuration it is a small slice of total memory. The first step in judging whether a new memory-sharing mechanism has an object is to subtract the file sharing that already exists, then ask how much of the anonymous remainder is duplicate, safe to share, and still there later.
只读共享 repo 之后,同 repo sandbox swarm 剩下的可去重内存只有 4% 到 10%,而且几乎全是各自的私有堆,不是文件。
coding agent 的执行形态正在从一个会话一个 sandbox,变成一个任务同时派出一群 sandbox。subagent 把子任务分给并行的执行环境,best-of-N 在同一个问题上同时试 N 种改法,强化学习的 rollout 在同一个仓库快照上并发跑成百上千条轨迹。它们有一个共同点:同一台宿主上的这些环境读的是同一个仓库、同一套依赖、同一个编译器与语言运行时。于是一个很自然的问题是,宿主内存是不是被大量内容相同的页各存了一份,单机能跑的并发数是不是被这些重复卡住。这篇讲我们怎么测这个问题,以及为什么测出来的答案让一个新的共享内存抽象失去了对象。
隔离环境里的一页内存从哪来
同一个仓库上的 sibling sandbox 里,文件页和匿名页的共享机制完全不同。
flowchart LR
Repo["共享 repo"] --> PC["宿主 page cache<br/>同 inode 一份"]
Tool["tsserver / pytest"] --> A1["A 的堆"]
Tool --> A2["B 的堆"]
Tool --> A3["C 的堆"]
PC --> S["文件页通常一份"]
A1 --> D["重复多在匿名堆"]
A2 --> D
A3 --> D
问题不应笼统地问 sandbox 是否重复占内存,而应问:在现成文件层共享已经生效之后,还剩多少内容相同的匿名页值得单独做跨 sandbox 去重。
先交代读懂后文需要的量级与机制。Linux 的物理内存以 4 KiB 页为单位,一个进程看到的内存分成两类。文件页是某个文件内容在内核 page cache 里的缓存,读文件、mmap 可执行文件与共享库都落在这里;同一个 inode 的同一偏移在宿主上只有一份物理页,无论多少进程在读。匿名页(anon)是没有文件背书的内存,进程的堆、栈、JIT 代码与语言运行时的对象都在这里,默认每个进程一份。
容器靠 overlayfs 共享文件页。overlayfs 把一个只读的 lower 目录与一个可写的 upper 目录叠成一个视图,读未修改的文件直接落到 lower 的同一个 inode,写的时候把整个文件 copy up 到 upper 私有化(Linux overlayfs)。所以同一镜像的一百个容器读同一个库文件,宿主上只有一份 page cache。代价是写放大按文件而不是按页,修改一个大文件的一个字节要复制整份。
虚拟机没有这个便利。microVM 的 guest 有自己的内核与自己的 page cache,它通过 virtio-blk 读一个块设备镜像时,内容先进宿主 page cache(如果宿主没有用 O_DIRECT),再被拷进 guest 内存里的 guest page cache,从宿主看后者是 hypervisor 进程的 anon 页。N 个 guest 读同一个镜像,宿主上就有 N 份 guest page cache。Firecracker 是面向 serverless 的极简 VMM,设备模型只有 virtio-net、virtio-blk、vsock 等少数几种(Agache et al. 2020)。解决 guest 侧重复的做法是 DAX(direct access):virtiofs 的 DAX 窗口或 virtio-pmem 设备把宿主文件直接映射进 guest 物理地址空间,guest 文件系统以 DAX 方式挂载时不再建自己的 page cache,读的就是宿主那一份页(virtio-fs)。Firecracker 从 1.14 开始支持 virtio-pmem,文档写明 backing file 以 MAP_SHARED 映射、可设只读、guest 用 DAX 挂载时绕过 guest page cache,同时提醒多个 VM 共用同一 backing file 会构成跨 VM 的侧信道(Firecracker pmem)。
还有两类共享发生在别的时间点。一类是出生即共享:fork 让子进程在写之前与父进程共享全部页,zygote 与 fork server 先把运行时预热好再 fork 给每个任务;Firecracker 快照恢复用 MAP_PRIVATE 映射内存文件,从同一快照恢复的多个 VM 在分歧前共享干净页。REAP 记录恢复后真正触及的工作集并提前整批装入(Ustiugov et al. 2021),FaaSnap 进一步按页访问顺序组织快照(Ao et al. 2022),Catalyzer 用 sfork 从一个已初始化的模板 sandbox 派生新实例(Du et al. 2020)。它们主要优化冷启动时间,顺带让出生时刻的内存共享。另一类是事后合并:KSM 周期扫描标了 MADV_MERGEABLE 的匿名区域,把内容相同的页合并成一份写时复制页,代价是扫描 CPU、合并前的峰值与已知的侧信道(Linux KSM);Medes 把这一思路做到 serverless 平台,跨 sandbox 按内容去重暖实例的内存(Saxena et al. 2022)。
这些机制合起来覆盖了三段:出生时共享(fork、快照、zygote)、运行中读共享(overlayfs、DAX)、事后合并(KSM、Medes)。后文的每个比例都要跟这个组合比,而不是跟什么都不共享比。
量级参照:一个在中等 TypeScript 仓库上常驻的 tsserver 大约 0.5 GB,其中九成以上是 anon;一个 Node 20 加 Python 的最小工具链镜像约 0.2 GB 文件,真正被读进 page cache 的不到 0.15 GB。也就是说,一个 sibling 的常驻集里,文件那部分本来就小。
为什么会怀疑存在大量重复页
我们的出发点是一个具体的抽象:宿主维护一个按内容寻址的只读页对象命名空间,repo、依赖与工具链的每个版本是一组不可变对象,所有 sibling 从出生起映射同一份,写入只影响写入者,版本切换就是换一组对象。它的直觉是 sibling 之间读集高度重合,swarm 越大,不共享的浪费越线性增长。
在进一步实现前,把这个想法对着现成机制核对,已经看到了风险:文件页那一侧,overlayfs、DAX 与 KSM 组合起来已经表达了出生即共享、写时私有化的定义性能力,剩下的只有可变 anon 堆里碰巧相同的页,而那不是不可变对象。我们当时用上一轮对开发工具进程的内存普查做了一个量级估算,那次普查测得 tsserver、常驻 pytest worker 与 javac daemon 的常驻集里 anon 占 91% 到 96%。按这个比例,32 个 sibling 时逐页去重 oracle 相对 DAX 配置再省 6.6% 到 9.4%,低于我们设的 20% 的预设收益阈值。
但这是估算,前提是 swarm 里的 anon 比例与单进程同量级,而且 guest 内没有大块 DAX 覆盖不到的只读派生数据。这两条都可能不对:一群 sibling 同时读同一仓库时,read buffer、解析出来的源码字符串与编译产物可能在各自堆里各存一份相同内容,那是不可变内容进了 anon,文件共享机制拿不到,正是新抽象可能有价值的地方。所以我们决定实测。
TL;DR
subagent、best-of-N 和 RL rollout 会在同一个仓库上同时跑几个到几十个 sandbox。它们读的是同一份 repo、同一套依赖和同一套工具链,所以很容易以为宿主里有大块相同的页各存了一份,值得做一个出生即共享、按内容寻址的只读命名空间。
我们在一台 112 核的机器上,让 4 到 16 个 sibling 对同一个 TypeScript 仓库跑 10 分钟。负载是常驻的 tsserver、grep、tsc、pytest,以及各自改自己的文件。逐页内容哈希比较了四种配置,对照是全 swarm 去重后的 oracle:宿主进程、共享只读 repo 的容器、每个 sibling 自己拷一份 repo 的容器、不做共享的 Firecracker。
只要 repo 和工具链以只读层进来,10 分钟后还没被共享的重复只占 swarm 内存的 4% 到 10%。其中大约 96% 是每个 sibling 自己的匿名堆。文件页已经只有一份。堆里的重复又有 52% 到 92% 在进程刚起来时就在了,是同一套工具按同样方式启动的结果。只有完全不共享的 Firecracker 留下 37% 到 42% 的重复,主要是每个 guest 自己的 page cache 和内核页。virtio-pmem 加 DAX 就是用来去掉这一块的。强配置下,这个命名空间要管的对象几乎已经不在了。
普查怎么做
图 1 是被测的两种极端。左边是容器或宿主进程读同一个只读层,repo 与工具链的页在宿主上只有一份,每个 sibling 私有的是自己工具进程的堆。右边是 Firecracker 通过只读 virtio-blk 挂同一个 rootfs 镜像,宿主 page cache 里镜像只有一份,但每个 guest 的内存里还有自己的 page cache 副本与自己的内核。
硬件是一台 112 核、251 GB 内存的 x86 服务器,宿主内核 5.15,只用 CPU 与 KVM。负载是一个固定 commit 的 date-fns 仓库(TypeScript,TypeScript 5.9 编译器)加一个 Python 小库 more-itertools。每个 sibling 跑 10 分钟:一个常驻 tsserver 打开 200 个源文件并做诊断、quickinfo、查引用与补全,每轮 grep 全仓库、改一个自己名下的文件并推给 tsserver,每三轮跑一次 tsc 全量类型检查,每两轮跑一次 pytest。每个 sibling 的文件选择由自己的编号作种子,所以它们做的事相似但不相同。
四个配置。宿主进程臂里每个 sibling 是一组宿主进程,通过各自的 overlayfs 读同一个工作目录。共享只读容器臂模拟 Docker:独立的 mount、pid、ipc、uts 命名空间,chroot 进一个由同一镜像 lower 层叠出来的根,repo 同样以只读 lower 加私有 upper 挂入。拷贝容器臂与之相同,只是每个 sibling 拷一份自己的 repo。Firecracker 臂用 1.10.1,每个 VM 2 个 vCPU、1.5 GiB 内存,只读挂同一个 ext4 镜像,guest 里跑同样的 sibling 脚本。本机的 Firecracker 版本没有 virtio-pmem,所以 DAX 臂没有测,这是本文最重要的一个未测项。
普查在第 15 秒、1 分钟、5 分钟、10 分钟四个时刻做。每次先用 cgroup freeze 冻住所有 sibling,读每个进程的 pagemap 取出所有在内存里的匿名页,按物理页号去重得到实际占用,再逐页算内容哈希;对 swarm 能看到的所有文件用 mincore 找出在 page cache 里的页,按 inode 计实际占用、按内容哈希算 oracle。实际占用是文件页加去重后的匿名页,oracle 是全 swarm 文件与匿名页并集里内容不同的页数,两者之差就是内容相同却没被共享的重复。重复再按来源分三类:不同 inode 里内容相同的文件页;内容等于某个不可变文件块的匿名页(guest page cache、read buffer、拷贝进堆的文件内容);其余匿名重复,即各 sibling 私有堆里碰巧相同的页,这一类再按内容是否已在第 15 秒的快照里出现拆开。Firecracker 臂里 guest 的全部内存都是 VMM 进程的匿名页,所以 guest page cache 落在第二类。每格跑一次。
N 取 4、8、16。共享只读容器臂与 Firecracker 臂的 N=16 以及所有臂的 N=32 没有跑完:前两个臂在 N=4 与 N=8 时结论已经远离阈值,我们在那时停了探针。
测得了什么
图 2 左边是 10 分钟时刻每个配置、每个 N 下没被共享的重复内存,柱高是 GB,柱顶数字是它占 swarm 实际内存的比例,颜色是来源。右边是每个配置在最大 N 下每个 sibling 平均的实际占用(浅灰)与 oracle(深灰)。
宿主进程与共享只读容器两个配置,在 N=4 到 16、1 到 10 分钟的全部 15 个采样点上,重复都在 3.8% 到 9.8% 之间,10 分钟时是 5.6% 到 8.2%。颜色几乎全是灰与绿,也就是私有堆:宿主臂 96% 到 97%,容器臂 90% 到 92%。文件页一侧的重复是 0% 到 2%,每个 sibling 读进来的 repo 与工具链在宿主上确实只有一份,swarm 的文件页总量在 N=4 到 16 之间一直停在约 0.15 GB(宿主)与 0.11 GB(容器),不随 N 增长。swarm 占用随 N 线性增长的是 anon,N=16 的宿主臂 6.76 GB 里有 6.61 GB 是 anon。
我们事先最担心的那一类,不可变文件内容进了各自的堆,在这两个配置里只占重复的 3% 到 9%,也就是 swarm 内存的千分之几。tsserver 与 node 把源码读进来之后,存下来的是 AST、符号表与类型信息,而不是原文的 4 KiB 页。
拷贝 repo 的容器把重复推到 16% 到 21%,其中约 70% 是文件页:每个 sibling 有自己的 repo inode,内容相同的源文件在 page cache 里各存一份。这个配置正是 overlayfs lower 层要避免的,它说明问题本身是真实的,只是已经有了标准解法。
无共享 Firecracker 的重复最大,10 分钟时 38.5% 到 38.8%,1 分钟时最高 41.5%。来源里 17% 到 29% 是文件页(宿主 page cache 里镜像与 guest 缓存内容相同的部分),39% 到 47% 是内容等于镜像块的 guest 内存,即 guest page cache,其余是 guest 里私有堆与内核的相同页。右图里每个 VM 平均 0.62 GB,oracle 只要 0.38 GB。这是唯一一个超过 20% 阈值的配置,而超出的部分主体是 guest page cache,virtio-pmem 加 DAX 让 guest 直接映射宿主的那一份,正好移走这一块。我们没有实测 DAX 臂,所以不能给出它之后剩多少,但按共享只读容器臂的构成,剩下的主要是堆与 guest 内核。
堆里的重复是创建时带来的
图 2 里每根柱子的灰色斜线部分,是在第 15 秒快照里就已存在的私有堆重复。它在全部 24 个 1 分钟以后的采样点里占私有堆重复的 52% 到 92%,宿主与共享只读容器臂在 61% 到 92%。换句话说,sibling 之间相同的那点堆,大多不是运行中偶然长成一样的,而是同一个 node 与 tsserver 二进制、同一个 Python 解释器在相同输入上做完启动与初始化之后的产物。按内容相同的页被多少个 sibling 同时持有统计,在三个非 VM 配置里,约 3200 到 3300 页(约 13 MB)在每个 N 下都被全部 sibling 同时持有,数目几乎不随 N 变化,是每份工具启动都会生成的那部分。
这一块已经有主人。zygote、fork server 与快照恢复在出生时刻让这些页天然只有一份,运行后被写才私有化。我们在上一轮对同类工具的普查里看到,重放同一请求序列后只有 3% 到 12% 的页与原进程内容相同,GC、分配顺序与哈希种子让两个回答完全一致的堆在字节上迅速分开;这一轮 swarm 里运行中新长出的相同页不到 swarm 内存的 2%,与之一致。
为什么这个判断不成立
假设要成立需要两件事同时为真:在强共享配置下还有大块内容相同的页没被共享,并且它们是不可变内容,适合做成出生即共享的对象。实测里第一件在强配置下不成立,剩余只有 4% 到 10%,离 20% 的阈值一半都不到;第二件也不成立,剩余的九成以上是私有可变堆,按定义不在不可变对象命名空间的管辖之内,其中出生时就存在的那部分又已被 fork 与快照拿走。唯一有大空间的无共享 microVM 配置,空间来自 guest page cache,DAX 与 virtio-pmem 已经是它的标准解法。决定性的是共享只读容器臂那一行,它是 Docker 默认行为,不需要任何新机制。
还有几处没覆盖。每个格子只跑了一次。共享只读容器和 Firecracker 没有 N=16 和 N=32。宿主这边,重复比例从 N=4 到 16 没有上升,依次是 5.6%、8.2%、5.9%。我们据此外推 N=32 不会越过阈值,这是外推。负载主要是中等仓库上的只读分析。数 GB 的 node_modules 或大型 monorepo 会把文件页放大,但放大的正是已经被共享的那一块,方向对结论有利。长于 10 分钟的会话、同一台机器上的多个仓库,以及进了私有堆的只读派生产物,都没有覆盖。索引、编译缓存和向量库属于最后这一类。它们可能让不可变内容进匿名页的比例上升,也是这篇结论最可能失效的地方。
共享机会真正剩在哪里
同一个 repo 上的 sandbox swarm,只要 repo 和工具链走只读层(宿主 overlayfs,或容器镜像的 lower 层),10 分钟时还没被共享的重复只占实际内存的 4% 到 10%。其中九成以上是每个 sibling 自己的匿名堆。这是 4 到 16 个 sibling、中等 TypeScript 与 Python 仓库、语言服务器常驻再加测试和构建、4 KiB 页、每个格子只跑一次。
随 sibling 数量线性涨的是匿名页。文件页的总量几乎不随数量变。开发工具的常驻内存本来就以堆为主,所以 swarm 的内存问题是每个 sibling 自己的工具进程堆,不是同一份文件存了多份。条件与上一条相同。
私有堆里、跨 sibling 内容相同的页,有 52% 到 92% 在进程刚起来时就已经在了。它们来自同一套工具按同样的方式启动。只读共享打开之后,跑起来才新长出来的相同页,不到 swarm 内存的 2%。要省这一块,应该在出生时共享,例如 zygote、fork server 或快照恢复,而不是在运行中按内容寻址,或者事后再扫描。前提是各 sibling 的工具二进制和输入相同。
完全不做共享的 microVM 群会留下 37% 到 42% 的重复。主体是每个 guest 自己的 page cache,以及内容相同的内核页。virtiofs DAX 和 virtio-pmem 就是用来去掉 guest page cache 的。估 microVM 群还能共享多少内存时,对照里必须有 DAX。否则会把 DAX 已经能拿到的收益,记到新机制头上。DAX 之后还剩多少,这次没有测。
一个共享机制能省的上限,是现成最强共享已经做完之后,剩下的重复里、刚好属于它要管的那一类。这次普查把剩余拆成四类:文件、进了堆的不可变内容、出生时就已经在的堆、跑起来之后才长出来的堆。任何声称能跨 sandbox 去重的方案,都可以先对一下自己管的是哪一类,以及那一类有多大。
先扣掉已经存在的文件共享
同 repo 的 sandbox swarm 看起来天然会复制很多东西,但 Linux page cache、overlayfs 和共享只读层已经提前拿走了最容易共享的那一部分。真正剩下的重复主要在私有匿名堆,而且在强配置下只占总内存的一小块。判断新的内存共享机制有没有对象,第一步应该先把现成文件共享扣掉,再看匿名页里还剩多少可重复、可安全共享、且会持续存在的内容。