A killed process writes no profile
Coverage went from 53% to 87% and every file crossed 80%. Three things cost a day each: a SIGKILL that silently erased a whole daemon's coverage, a test that passed by skipping, and a read with a deadline it never checked.
On this page · 4 sections
Seven tools shipped in a week - intent search, program slicing, shadow runs, a structural codemod, a fixture generator, a signature change and a cross-language rename. Each had tests, each was reviewed, each had a pull request with transcripts. Then somebody asked what the coverage actually was, and the honest answer took a while to produce and was worse than expected:
TOTAL 40996 19161 53.26%
Fifty-three percent, and useless besides: a workspace total lets one file at 0% hide behind a dozen at 90%. The question worth answering is not “what is the number” but “which file is nobody testing”, and no tool we had answered that.
This is what it took to get every file over 80%, and the three things that cost a day each on the way. Not one of them was a bug in the logic the tests were aimed at; two turned out to be bugs in the daemon, and the third was in the test harness.
What the existing tests were not testing
The unit tests covered parsers - case conversion, argument-spec parsing, splitting a parameter list without splitting the types inside it - and not one bug had ever been there. Every real bug lived in the composition: which analyzer answer gets merged with which, in what order, and what is reported when two disagree. That needs a gateway, and a gateway needs an analyzer, so the tests stopped at the edge of it.
Except it does not need an analyzer. It needs answers of the right shape. prod-code-testkit
is a real TCP server that speaks the real wire protocol - accepts the pre-flight sync, answers
the handshake, replies to initialize itself - and hands every other LSP request to a closure
the test provides:
let workspace = Workspace::new(&[("src/lib.rs", "pub fn a() {}\n")]);
let gateway = ScriptedGateway::start(|method, params| match method {
"textDocument/rename" if uri_of(params).ends_with("lib.rs") => {
// rust-analyzer answers a rename with the file's whole new text …
answers::whole_file(&path, OLD, NEW)
}
"textDocument/rename" => {
// … while gopls answers with one edit per occurrence.
answers::ranged(&go_path, &[(4, 2, 7, "TradeID")])
}
"textDocument/diagnostic" => answers::no_diagnostics(),
_ => serde_json::Value::Null,
})
.await;
Those two shapes in one test are not a detail: merging them is what turned a Rust file into duplicated fragments in the cross-language rename, and the test that scripts both is the one that holds the fix. Nine such tests run in 0.08 seconds with no analyzer, no node and no network.
scripts/coverage.py reports every file, sorted worst first, and fails on the files themselves:
$ prod-code exec --timeout-secs 1800 --no-pull -- python3 scripts/coverage.py --min 80
! crates/prod-code-client/src/main.rs 0.00% 0/3241
! crates/prod-code-gateway/src/backend.rs 7.97% 24/301
! crates/prod-code-gateway/src/main.rs 15.66% 908/5798
…
27 file(s) under 80%
and the same command is what says the work is done:
$ prod-code exec --timeout-secs 1800 --no-pull -- python3 scripts/coverage.py --min 80
crates/prod-code-gateway/src/lib.rs 80.46% 4675/5810
crates/prod-code-engine-generic/src/lib.rs 82.99% 732/882
crates/prod-code-client/src/main.rs 83.25% 2698/3241
…
TOTAL 87.26% 37022/42429
every file is at or above 80% of regions
The first line of that first report is the one that matters. The CLI - the thing every human and every script actually types - had 3241 regions and not one had ever run in a test. That is also why the gate has no exemption list: an exemption is a file nobody has to think about again, and this was exactly the file somebody should have been thinking about.
Three traps, none of them in the code under test
A killed process writes no profile. The gateway is a daemon. To test it for real you start the real binary, drive it, and stop it.
The test did that, coverage was collected, and the gateway’s main.rs moved from 15.66% to …
15.66%.
Nothing. Eleven tests exercising every dispatch arm, and not a single region attributed.
The reason is one line in the test’s cleanup:
impl Drop for Gateway {
fn drop(&mut self) {
let _ = self.child.kill(); // SIGKILL
let _ = self.child.wait();
}
}
The coverage runtime writes its .profraw at exit, through a handler registered at
startup. A process killed with SIGKILL runs nothing at exit. Every byte of that daemon’s
execution was measured and then thrown away, and the report showed exactly what it shows for
code that never ran: 15.66%, with no indication that anything had been lost.
Sending SIGTERM instead changes nothing on its own: the default action for SIGTERM also
terminates the process without running anything at exit. The profile appears only once the
daemon handles the signal and returns from main normally - and the daemon had no handler,
which turned out to be the more interesting problem.
Every supervisor and every deploy script already sends SIGTERM and then waits. launchd does, systemd does, our own deploy does. In-flight queries were being cut mid-answer in production, and had been since the first node was deployed. So the fix is not in the test:
// A gateway is stopped by its supervisor and by a deploy script, both of which send SIGTERM
// and then wait. Without a handler the process dies where it stands: in-flight queries are
// cut, and nothing that runs at exit runs.
let mut terminate = signal(SignalKind::terminate())?;
loop {
let (socket, addr) = tokio::select! {
accepted = listener.accept() => accepted?,
_ = terminate.recv() => return Ok(()),
};
…
}
With the handler in place and the test sending SIGTERM, that file went from 15.66% to 51.45% in one run. A coverage tool found a production bug by reporting a number that could not be true.
main - which needs both a signal that can be handled and a handler for it.A test that passes by skipping is not a test. Some tests need a real language server. gopls is installed on the build node, so this looked reasonable:
if which("gopls").is_none() {
eprintln!("skipping: gopls is not on PATH");
return;
}
It passed. It also finished in 0.06 seconds, which is not enough time to start gopls, let alone ask it anything.
which looked at the PATH of the test process. The daemon puts the user’s toolchain
directories first on PATH itself, at startup, which is why gopls works in production and why
nothing else had ever noticed. The test was deciding gopls was absent on a machine that has it,
and reporting that as success.
The fix is to use the same policy the daemon uses, and the lesson is the tell: a test that does real work and finishes instantly is not fast, it is lying. Coverage confirmed it - the file that test exists to exercise went from 7.97% to 83.72% once it stopped skipping.
A deadline checked between reads is not a deadline. The harness first picked a free port by
binding a socket and dropping it, and eleven gateways starting at once found the window in
between. The fix was to let the daemon bind port 0 and report the address it got - the second
production change below. With that in place, the suite hung.
Not slowly - completely. Eleven gateway processes alive, zero CPU time each, for half an
hour, while the test binary’s main thread sat in futex_wait_queue, waiting on test threads that
were themselves waiting on a read that would never return.
let deadline = Instant::now() + Duration::from_secs(60);
while Instant::now() < deadline {
reader.read_line(&mut line)?; // ← never returns
…
}
read_line on a pipe that stays silent blocks forever. The deadline is checked between reads,
so if the first read never returns, the deadline is decoration. And the pipe was silent for a
second reason: tracing_subscriber::fmt() writes to stdout, and the harness was reading
stderr.
Two small mistakes, and the symptom of both was a suite that looked like it was working hard. Load average said otherwise; zero CPU across eleven processes is not slow, it is stopped.
The read is on its own thread now, with the answer on a channel and recv_timeout - a deadline
that a blocked read cannot outlive.
What the slow suite buys
Eleven tests start the daemon - the prod-code-server binary the prod-code-gateway crate
builds - as a child process with its own storage, and drive it through the same client an agent
uses: a Rust checkout, a Go checkout, a TypeScript checkout,
a Python checkout, a C++ checkout, two gateways gossiping and placing work on each other, an
idle workspace being evicted and answering again afterwards, and the write path through a
rename. Four minutes, and they are the only four minutes in the suite that prove a workspace
loads at all.
Everything else - 400-odd tests - runs in under a minute against scripted answers.
| file | before | after |
|---|---|---|
prod-code-client/src/main.rs |
0.00% | 83.25% |
gateway/src/lib.rs (was main.rs) |
15.66% | 80.46% |
gateway/src/main.rs (15 lines, after the split) |
- | 100% |
gateway/src/backend.rs |
7.97% | 83.72% |
engine-generic/src/lib.rs |
21.25% | 82.99% |
mcp/src/tools.rs |
27.79% | 89.89% |
engine-rust/src/lib.rs |
75.40% | 90.05% |
| the workspace | 53.26% | 87.26% |
Two production changes fell out of the work, beyond the SIGTERM handler: the gateway crate became a library with a thin binary on top, because a test cannot so much as name a function inside a binary crate; and the daemon now logs the address it actually bound rather than the one it was asked for, which is the only way to learn the port when you ask for zero.
What the number still does not say
The scripted answers are shapes observed from real engines, not the engines themselves. A test like that cannot notice rust-analyzer changing how it answers a rename - only the live suite can, and only for the paths it walks.
And 87% is a measure of coverage regions that ran, not of assertions that mean anything. The guard against coverage theatre is not a percentage: it is that every test here asserts an observable outcome - a rendered string, a file’s content, an exit code, the number of calls a script received - and that both regression tests were checked by breaking the code on purpose and watching them go red. A test that cannot fail proves nothing, however much of the file it colours green.
Cite this article
Alexander Panasenko (2026-10-02). A killed process writes no profile. https://prod.codes/blog/a-killed-process-writes-no-profile/