design note · · 4 min

Language detection: what sixty test cases prove

prod-code v0.3.19 expands workspace detection, but a detected language is not proof of semantic support. The marker matrix and release artifacts show what the version provides.

On this page · 4 sections
  1. Extensions are not languages
  2. The verification matrix
  3. Sixty-three tools over one socket
  4. Distribution and the rule

An autonomous coding agent does not work in one language. In the morning it refactors a Rust gateway; by afternoon it touches Bazel Starlark rules, updates a Terragrunt configuration, adjusts a Typst documentation template, generates a WebAssembly text module, and edits a SystemVerilog testbench. Maintaining those toolchains on every agent host duplicates setup and caches. We put analyzer processes and builds on remote nodes so that clients can query code without installing each compiler locally. “Zero compilers” describes that client boundary; the nodes still need the relevant tools.

Text search alone cannot distinguish an identifier from a comment, cannot follow cross-file definitions through renamed imports, and cannot tell whether a type change broke callers. Code intelligence must remain semantic, but the toolchains belong on the cluster.

The v0.3.19 changelog describes an expansion from fifty to sixty languages. Its annotated tag resolves to commit d10fbd57188cb6c8b2ad3826b60c4f37862f7e9b; all code references below pin that snapshot. The useful distinction is between recognizing a workspace and having an analyzer capable of answering for it.

Extensions are not languages

The first mistake in multi-language tooling is trusting a file extension. A file ending in .v might be a hardware module in Verilog, or it might be source code for the V programming language. A file ending in .hcl could be a Terraform module, a Terragrunt orchestration root, or a Packer specification. A Makefile might drive a C project, or it might simply contain shorthand targets for a Python virtualenv or Go build.

A wrong engine choice can send a file to a server that does not understand it. Detection needs to account for project context before semantic requests are routed.

In detect.rs, prod-code uses an ordered cascade of manifest and source-file checks. This excerpt is abridged; the omitted branches also affect precedence:

// Abridged from detect_engine; intermediate branches omitted.
pub fn detect_engine(root: &Path) -> EngineKind {
    for marker in RUST_MARKERS {
        if root.join(marker).exists() { return EngineKind::Rust; }
    }
    for marker in GO_MARKERS {
        if root.join(marker).exists() { return EngineKind::Go; }
    }
    if has_swift_project(root) { return EngineKind::Swift; }
    // ...
    if has_starlark_project(root) { return EngineKind::Starlark; }
    if has_hcl_project(root) { return EngineKind::Hcl; }
    if has_systemverilog_project(root) { return EngineKind::SystemVerilog; }
    if has_vhdl_project(root) { return EngineKind::Vhdl; }
    EngineKind::Generic
}

Manifest precedence matters:

  1. Go and TypeScript manifests outrank Make detection: A directory containing package.json and a Makefile selects TypeScript before the Make-based C/C++ check. C/C++ Make detection requires both a Makefile and actual C source files (.c, .cc, .cpp) at the root or under src/.
  2. Disambiguating dialect siblings: Terraform detection recognizes .tf files as well as markers such as main.tf and versions.tf. A standalone terragrunt.hcl or .tflint.hcl is detected as HCL rather than Terraform. That classification alone does not start an HCL language server.
  3. Hardware description languages: SystemVerilog (.sv, .svh) and VHDL (.vhd, .vhdl) inspect hardware manifest markers (verilator.f, vunit.py) separately from V projects (v.mod, .v). A bare .v suffix remains ambiguous; these checks do not prove every mixed layout is resolved.

The verification matrix

The detection matrix contains sixty cases: 59 named language cases and one Generic fallback. It checks primary selection and whether the all-engines result contains the expected kind. It does not launch servers, request definitions, validate refactorings, or compile those projects.

This is an illustrative abridgment of selected cases and assertions, not the complete test. The Case declaration, other entries and fixture-writing loop are omitted:

#[test]
fn test_universal_language_detection_matrix_all_60_languages() {
    // Case declaration omitted.
    let matrix = [
        Case { lang: "Rust", files: &[("Cargo.toml", "[workspace]")], kind: EngineKind::Rust },
        Case { lang: "Starlark", files: &[("BUILD.bazel", "load(':rules.bzl', 'rule')")], kind: EngineKind::Starlark },
        Case { lang: "Wat", files: &[("main.wat", "(module (func (export \"run\")))")], kind: EngineKind::Wat },
        Case { lang: "Generic", files: &[], kind: EngineKind::Generic },
        // Other cases omitted.
    ];
    // Fixture setup omitted. Each case checks:
    // assert_eq!(detected, case.kind, ...);
    // assert!(all.contains(&case.kind), ...);
}

detect_all_engines evaluates markers in the supplied directory; this test is not a proof of recursive discovery throughout an arbitrary monorepo. Nested project routing is a separate mechanism.

The support boundary is visible in workspace engine loading and backend spawning. At this tag, the ten new detection kinds have no corresponding managed server branches there. Existing generic adapters also require their server binaries to be installed on the node. The matrix establishes recognition, not sixty working semantic backends.

Sixty-three tools over one socket

Once an engine is detected, where does execution happen?

The MCP tool catalog has 63 entries at this tag. That is the number of exposed operations, not a promise that each operation works for every detected language. A refactoring still depends on its implementation, project shape and the relevant analyzer.

The protocol transport supports TCP, Unix sockets under cfg(unix), and a Windows named-pipe client under cfg(windows). The added pipe connector opens an existing pipe and optionally sends an auth frame. It does not establish a shipped Windows gateway listener or a complete local agent bridge.

Remote execution moves analyzer indexing and compiler work to the node. The client still scans, hashes and synchronizes source files, translates paths and processes responses. We have no measurement here that supports zero local CPU cost or a general sub-second query guarantee.

Distribution and the rule

At publication, the v0.3.19 release assets include macOS .pkg and .dmg files, amd64 and arm64 .deb files, an Apple Silicon client binary, install.sh, and SHA256SUMS. Those artifacts exist, but their names do not establish signing, service installation or fleet parity. Inspection of the macOS package reports no signature; we cannot describe it as a signed installer.

The installer script maps OS and architecture to an asset name, downloads it, and moves it into a writable binary directory. It does not check SHA256SUMS or authenticate a release signature; its macOS signing step is ad hoc. The release lists only one standalone platform binary, so platform detection in the script is not proof that every detected target has a downloadable binary. The pinned CLI also has no package verify or package sync command.

The Zed extension does provide a concrete editor route: it invokes the local client with lsp --language <language>. No caching latency number follows from that wiring.

Our rule: test detection separately from semantic behavior, confirm that the chosen analyzer is installed and wired, and verify the actual release artifact before promising a working path.

Cite this article
Citation
Alexander Panasenko (2026-10-01). Language detection: what sixty test cases prove. https://prod.codes/blog/sixty-languages-zero-compilers/