design note · · 6 min

Stop letting agents test in your working tree

Every agent harness runs the same loop: write a candidate fix into your checkout, run the tests, revert, try the next one. We moved that loop into overlay shadows of the mirrored workspace, mounted at the workspace's own path, and it got faster, not just cleaner.

On this page · 5 sections
  1. The loop everyone accepts
  2. The same path, a different tree
  3. What a run looks like
  4. What it does not do yet
  5. Footnotes

Every coding agent you run today tests its ideas in your working tree. Claude Code, Codex, Gemini, the harness you wrote yourself: write a candidate fix, run the tests, read the failure, revert, write the next candidate. The loop is so normal that nobody calls it what it is, a scratchpad in the one directory that is supposed to hold your intent. The cheap half of that loop, asking the analyzer whether a patch even compiles, is a separate post; this one is about the half that has to run the tests. Between attempts the tree holds a fix that is not going to survive, and whatever reads it in that window, the resident session, a second worker, a file watcher, a pre-commit hook, reads something nobody meant.

We think this is the wrong place for the loop, and not by a little. The loop is also serial, because there is one tree: three candidates cost three test runs plus two reverts, and the agent cannot ask the question it actually has, which of these three passes.

The loop everyone accepts

The usual defences do not hold up. git stash between attempts still writes every candidate into the tree and still runs them one at a time; it is the same loop with more git. A worktree per attempt keeps the tree clean and throws away the build cache: a fresh checkout of our gateway repository needs 8.7 s just to build the dependencies of its smallest crate before cargo test can start, against 1.3 s for the test itself. Running it all on the laptop is the option most people have settled on, and it is why your machine is unusable while the agent works.

We run code intelligence for our agents off the laptop. prod-code1 mirrors a checkout to a Linux node on the LAN, keeps the analyzer warm there (rust-analyzer in process; gopls, clangd, the TypeScript server and pyright as child processes) and runs builds and tests in the mirror; agents reach it through MCP tools. So the question was how to run the same command three times against three versions of the mirror without three mirrors, and without three cold builds.

The same path, a different tree

A copy per candidate fails on the first cargo test: the copy has no target/, and copying target/ along does not help, because cargo’s fingerprints and the -L dependency=… paths it passes to rustc carry the workspace path2. A candidate that lives at a different path is a cold build.

So the candidate lives at the same path. Every hypothesis runs under unshare -Urm: a new user namespace, in which the gateway’s uid is root, and a new mount namespace, in which it mounts an overlay over the workspace directory itself:

mount -t overlay overlay \
  -o "lowerdir=$SHADOW_LOWER,upperdir=$SHADOW_UPPER,workdir=$SHADOW_WORK" \
  "$SHADOW_LOWER"

Inside that namespace /home/…/workspaces/repo is still the workspace: the same absolute path, the same Cargo.lock, the same warm target/, the same ~/.cargo registry. Reads fall through to the real directory; writes go to the hypothesis’s upper directory. The files that make the hypothesis different are written into the upper directory before the mount, so once it is up the tree is the workspace with those files replaced. Deletions are rm inside the mount, which overlayfs records as whiteouts. The whole runner is this script:

#!/bin/sh
# Generated by prod-code for one hypothesis of a shadow run.
set -e
mount -t overlay overlay -o "lowerdir=$SHADOW_LOWER,upperdir=$SHADOW_UPPER,workdir=$SHADOW_WORK" "$SHADOW_LOWER"
if [ -s "$SHADOW_DELETE" ]; then
  while IFS= read -r p; do rm -rf "$SHADOW_LOWER/$p"; done < "$SHADOW_DELETE"
fi
cd "$SHADOW_LOWER/$SHADOW_SUBDIR"
exec "$@" 2>&1

Nothing outside the namespace sees the mount. The resident rust-analyzer keeps its view of the real directory, another agent’s code_test runs against the real directory, and when the process tree exits the mount goes with it. What remains is the upper directory, the recompiled crate and its test binary, 47 to 80 MB for a small crate, which the gateway deletes after the run.

Three hypotheses, each with its own upper layer, all reading through one shared workspace copy mounted at the same path
Reads fall through to the shared copy, so the build cache stays valid; writes land in a layer that is deleted when the run ends.

Measured on a 32-core node (Ubuntu 22.04, ext4) with a warm workspace: a no-op cargo build inside an overlay reports Finished in 0.86 s, so the fingerprints survive the mount. Appending a function to one crate and building it takes 0.24 s, that crate only.

What a run looks like

A hypothesis is a name and the complete contents of the files it changes. The CLI takes them from a JSON spec; the MCP tool code_shadow_run takes the same shape inline:

{
  "hypotheses": [
    { "name": "baseline" },
    { "name": "plus",  "edits": [ { "path": "src/lib.rs", "file": "hyp/plus.rs" } ] },
    { "name": "mul",   "edits": [ { "path": "src/lib.rs", "file": "hyp/mul.rs" } ] },
    { "name": "clamp", "edits": [ { "path": "src/lib.rs", "file": "hyp/clamp.rs" } ] }
  ]
}
$ prod-code shadow-run spec.json -- cargo test
$ cargo test   (4 hypothesis(es), overlay mode, on …/workspaces/shadow-fixture)
  plus      exit 0     in    0.4s  2 passed, 0 failed  ~2 changed line(s)  <- winner
  clamp     exit 0     in    0.4s  2 passed, 0 failed  ~2 changed line(s)
  baseline  exit 101   in    0.0s  0 passed, 2 failed  ~0 changed line(s)
  mul       exit 101   in    0.4s  0 passed, 2 failed  ~2 changed line(s)
winner: plus
--- a/src/lib.rs
+++ b/src/lib.rs
@@ -1,6 +1,6 @@
 pub fn add(a: i32, b: i32) -> i32 {
-    a - b
+    a + b
 }

The fixture is a one-function crate whose add subtracts. baseline has no edits and answers the question that usually comes first: are the tests red right now? The gateway decides nothing; it returns exit code, duration and the output tail per hypothesis. The client parses test counts with the parsers it already had for code_test (cargo, go test -json, pytest, unittest, vitest, jest, bun, XCTest, ctest, meson) and sorts: passed before failed, fewer failing tests, more passing tests, a smaller diff, then the order given. Only the winner’s unified diff is printed; the output of every failing hypothesis follows it, tail only, so the agent sees the assertion that broke mul without asking again. apply: true writes the winner into the checkout. Nothing else ever touches the checkout, which is the point.

On this repository, three hypotheses of cargo test -p prod-code-protocol (baseline, an added passing test, an added failing test):

through the checkout, one at a time as shadows
three test runs 3 × (write, run, revert) 3.45 s server-side, in parallel
wall clock from the laptop — 2.0 s including the pre-flight sync3
the working tree meanwhile holds two fixes that will be thrown away untouched
left on the node afterwards — nothing

What it does not do yet

  • macOS has no overlayfs. A gateway without user namespaces runs hypotheses one after another in the node’s own copy of the workspace, restoring each file afterwards if its content is still what the hypothesis wrote. Your checkout is never touched in either mode, but on that path a crashed gateway can leave the node’s copy holding a hypothesis’s files. Our two Swift nodes are Macs and have not been updated to this build yet.
  • Ubuntu 24.04 ships with kernel.apparmor_restrict_unprivileged_userns=1; unshare -Urm fails with Operation not permitted and the node runs the same in-place mode until the sysctl is changed. The gateway probes this once at start and logs which mode it is in.
  • The lower directory is not frozen. A sync from the checkout that lands while a hypothesis runs changes files under a live overlay, which overlayfs promises nothing about.
  • Hypotheses do not persist between calls. Each call resends full file contents, and each upper directory is deleted after the run.
  • Cargo’s package-cache lock lives in ~/.cargo, outside the namespace, so parallel hypotheses queue on it for a moment; the report drops those Blocking waiting for file lock lines.

The rule is the same one that made us revert a shared target/ between git worktrees the day before, rather than let ten workers queue on one build lock: things that run at once do not share writable state. Your working tree holds what you meant. Everything an agent is not yet sure about runs somewhere else, and that somewhere else has to be as warm as the tree itself, or nobody will use it. A shadow shares the workspace’s files and caches by reading through them, writes into its own directory, and disappears. If your harness’s answer to this is git stash, it is testing in your tree with extra steps.

Footnotes

  1. github.com/alex09x/prod-code: one gateway per node, a checkout mirrored by content hash, one isolated workspace and analyzer per git worktree. ↩

  2. We learned this from sccache, installed as rustc-wrapper to warm new worktrees faster: it never hit across target directories, because the compiler command line it hashes contains the target path. ↩

  3. Measured on 2026-09-21 through the cluster’s placement, laptop to a 32-core Linux node, pre-flight sync included. ↩

Cite this article
Citation
Alexander Panasenko (2026-09-25). Stop letting agents test in your working tree. https://prod.codes/blog/stop-letting-agents-test-in-your-working-tree/