design note · · 6 min

Your comments are the search index

An agent that knows what it wants but not what it is called has nothing to ask. We gave it a question box, then measured what decides whether the answer is any good: not the ranking, the prose your team did or did not write above each function.

On this page · 4 sections
  1. There is no magic in it
  2. What the comment buys you
  3. What it is actually for
  4. What it does not do yet

The symbol index answers a question an agent rarely has. It knows every name in the workspace and will find pick_node the instant you type pick_node, which is useful exactly when you already know the answer. The question an agent actually arrives with is “where do we decide which node runs a workspace”, and nothing in the codebase is called that.

So it greps. Grep matches letters, returns lines rather than the declarations that own them, and gives an agent a hundred hits to read at the cost of its context window. Then it guesses a synonym and greps again.

We gave the gateway a question box instead. The gateway is the process on a LAN node that holds a mirror of the checkout and a warm analyzer for it; code_search is a tool it serves, ranking every declaration in the workspace against the words of a question. Building it took an afternoon. Measuring it taught us something about our own repositories that the feature itself did not.

There is no magic in it

It is worth being blunt, because a tool that implies more than it does gets trusted more than it should. Nothing in this path is a model. Four mechanical steps:

  1. the question is split into words, stopwords dropped, simple plurals folded: “where do we decide which node runs a workspace” becomes decide node run workspace;
  2. every declaration in the workspace was split the same way when it was indexed, across four fields: its name (close_open_orders → close open order), its container, its signature, and the doc-comment block directly above it;
  3. BM25 over those fields, the name weighted three times the doc, multiplied by a kind weight so a function outranks a field with the same name, and by a bonus for covering more of the question;
  4. sort, take ten.
The question \u201cwhere do we cancel an order\u201d becomes the words cancel and order; every declaration was split the same way into name, container, signature and doc comment, weighted three, one and a half, one point two and one; BM25 ranks them with tests excluded, and the winning hit was found through its comment rather than its name
The question and the code are reduced to the same thing: bags of words. The meaning comes from a human having written some of those words above the function.

The reason it reads as understanding is step 2’s last field. A doc comment is written in the same human language as the question, by someone who already did the work of saying what the code is for. The search adds your words to theirs. That is the whole trick, and it has a hard edge: ask in a language the code does not use and you get nothing, correctly.

$ prod-code search "otmena zayavki" --limit 1
no declaration matches `otmena zayavki` (7227 declarations in 476 files searched; the index
is lexical, so a question sharing no words with the code or its comments finds nothing)

An agent does not use the CLI. It calls the tool, and gets the same answer:

{ "name": "code_search",
  "arguments": { "query": "how do we decide which node runs a workspace", "limit": 1 } }
1 hit(s) for `how do we decide which node runs a workspace` in 22 ms (1012 declarations, 40 files)

 1. [function] pick_node  crates/prod-code-mcp/src/cluster.rs:107
    pub async fn pick_node(
    Chooses the gateway for `workspace_name` among `nodes`: the remembered placement when it
    is still one of the nodes, alive and able to serve `engine`, otherwise the quietest alive
    node that can serve it

Nothing in the name matches the question. The comment supplies the words the question used, decide and node and workspace, and that is what put it first.

Two corrections to that arithmetic came from reading real output rather than from writing tests, and both are in the code now.

A test outranks the thing it tests. Ask “where are hypotheses run in an overlay” and the best lexical match is overlay_hypotheses_run_in_parallel_and_leave_the_workspace_untouched, because a good test name repeats every word of the behaviour it covers. It is the right answer to a different question. Tests are now excluded unless the question mentions tests, spec, fixture or mock.

A scanner cannot tell a fixture from code. Our own search results contained a function that does not exist: a snippet inside a multi-line string literal in a test, indexed as if it were a declaration. Everything after a file’s test module begins is now treated as test material, which covers the fixtures without needing to parse string literals.

What the comment buys you

Here is the same repository, a trading system, asked two ways. The first query shares words with a function’s name; the second shares words only with a comment.

$ prod-code search "close open orders" --limit 1
 1. [function] HttpClient::close_all_open_orders  src/api2/bitmart/api.rs:158
    pub async fn close_all_open_orders(&mut self) -> Result<()>

$ prod-code search "batch cancellation critical for risk management" --limit 1
 1. [function] HttpClientOKX::close_open_orders  src/api2/okx2/api.rs:581
    pub async fn close_open_orders(&mut self, open_orders: Vec<OpenOrder>) -> ResultStd<()>
    HOT PATH: batch order cancellation - critical for risk management

The second hit is not findable by name. Whoever wrote that comment made a function reachable by intent, five years before anybody searched for it that way.

Which suggested a measurement we had not thought to take. We asked two or three intent questions of each of two repositories, counted how many of their top-level declarations carry a comment, and scored a question as answered when any of the first three hits was the code that implements the behaviour asked about. Our judgement, by hand, on a sample this small:

repository language declarations with a doc comment questions answered
a signals service Go 2464 1407 (57%) 2 of 3
a trading system Rust 5386 659 (12%) 1 of 2, and that one through a comment

The tool is identical in both. In this sample the repository with more comments answered more intent questions, which is a correlation on five questions, not a demonstration: it does not separate the doc field from the name, the container and the signature, and Go and Rust name things differently. What it does show concretely is the shape of the failures. On the 12% repository, “how do we handle a websocket disconnect” returns a URL helper and a metadata decoder, because those are the only declarations whose names carry those words; the one question that worked, “where do we cancel an open order”, landed on a function through its comment rather than its name. On the 57% repository, “how do we detect whale accumulation” lands on WhaleAlertDetector, doc comment and all.

We have been treating comments as a courtesy to the next human. They are now also the index that decides whether a machine can find the code at all.

What it is actually for

Finding the function is not the point. Getting to a correct edit without reading the repository is. Three calls, on a codebase the agent has never seen:

# 1. where is the thing I was asked about?
$ prod-code search "how do we decide which node runs a workspace" --limit 1
 1. [function] pick_node  crates/prod-code-mcp/src/cluster.rs:107

# 2. what does it depend on? (not: what file is it in)
$ prod-code slice pick_node --depth 1
slice of `pick_node`: 1 item(s), 214 bytes from 18314 bytes of source (99% smaller)
outside the workspace, not followed: Option, Result, SocketAddr, as_deref

# 3. is the patch I drafted even correct? nothing is written
$ prod-code validate crates/prod-code-mcp/src/cluster.rs --from draft.rs
crates/prod-code-mcp/src/cluster.rs: 1 error(s), 0 warning(s)
  error: expected Option<&Path>, found Option<PathBuf> [E0308] (cluster.rs:114:51)
[prod-code] analysed in 2.15s

Two hundred and fourteen bytes of code read instead of eighteen thousand, and the type error caught before a single byte reached the disk, let alone a compiler. The search is the first step of that chain and the cheapest one; it is worth building because of what it makes possible afterwards, not because ranking declarations is interesting.

What it does not do yet

  • No embeddings: a question sharing no words with the code or its comments finds nothing.
  • No stemming beyond a crude plural fold, so cancellation and cancel are different words.
  • The index is per workspace and lives in the gateway’s memory, so it is rebuilt on first use after a restart. On a 32-core Linux node, against a 1051-file Go and Swift repository whose index holds 26712 declarations, the first query took 568 ms because it built the index, and the next three queries took 33 ms each as the gateway timed them.
  • Declarations are found by a line-based scanner, not the analyzer. An exotic declaration shape costs a missing hit; a multi-line string containing declaration-shaped lines can produce a wrong one, which is why everything after a file’s test module is excluded, and why a fixture in a string outside a test module would still be indexed as code. code_symbols remains the authority on what a name resolves to.
  • Comments are indexed as written. A stale comment makes code findable under a description of what it used to do.

The rule we took from it: a doc comment is no longer documentation, it is an interface. Not to the next reader, to the next search. And if you run agents against a codebase where nobody writes them, the cheapest upgrade available to you is not a better model or a bigger context window. It is one sentence above each function saying what it is for.

Cite this article
Citation
Alexander Panasenko (2026-09-26). Your comments are the search index. https://prod.codes/blog/your-comments-are-the-search-index/