notes · · 3 min

Detekt Under the Microscope: What 67 Remote AST Tools Found Inside Kotlin's Static Code Analyzer

We evaluated prod-code AST tools against detekt/detekt: 132,028 lines of Kotlin across 1,105 files, 41-module Gradle architecture with 4 test cycles, structural search across 31 files, and cluster-verified AST refactoring.

On this page · 4 sections
  1. The 41-Module Hierarchy and Test-Harness Dependency Cycles
  2. Rule Specification Clones: Parameterized Linting Harnesses
  3. Structural Metavariables: Precondition Audits Across 31 Files
  4. Polyglot Kotlin AST Refactoring and Cluster Language Server Verification

Static code analysis engines occupy a privileged tier in software infrastructure. By inspecting abstract syntax trees, resolving types, and reporting code smells across millions of lines of source code, tools like Detekt (detekt/detekt) safeguard maintainability for Kotlin development teams worldwide. Operating on top of JetBrains’ Kotlin compiler Program Structure Interface (PSI) and the modern Analysis API, Detekt evaluates hundreds of specialized inspection rules covering complexity thresholds, coroutine anti-patterns, naming standards, and exception handling discipline.

Evaluating a static analyzer with automated code intelligence infrastructure introduces a compelling meta-analytical experiment: what happens when an engine designed to critique abstract syntax trees is itself examined by a suite of 67 remote AST tools?

Across its modular Gradle structure, Detekt contains 132,028 lines of Kotlin across 1,105 source files:

$ git ls-files '*.kt' | wc -l
1105
$ git ls-files | wc -l
1993
$ git ls-files '*.kt' | xargs wc -l | grep -v 'total$' | awk '{s+=$1} END {print s}'
132028

We deployed prod-code’s suite of 67 AST and code intelligence tools against Detekt HEAD on remote cluster nodes. By offloading every dependency parse, clone comparison, AST query, and compiler verification pass to the cluster, we maintained zero CPU overhead on the local workstation while extracting the deep structural mechanics of Kotlin’s primary linter.

The 41-Module Hierarchy and Test-Harness Dependency Cycles

Modern Kotlin tooling codebases divide functionality into targeted components: public rule interfaces, core execution engines, reporting formats, and rule bundles. Detekt structures its build across 41 subprojects managed through Gradle Kotlin DSL scripts (settings.gradle.kts).

Executing prod-code dependencies across the repository reconstructed the complete inter-module dependency topology:

$ prod-code dependencies
discovered 41 modules across Gradle workspace:
  • detekt-api, detekt-core, detekt-cli, detekt-parser, detekt-psi-utils
  • detekt-rules (and 12 specialized rule subprojects)
  • detekt-test, detekt-test-utils, detekt-test-junit, detekt-test-assertj
  • detekt-report-sarif, detekt-report-html, detekt-report-markdown
  ... (16 additional integration and tooling modules)

inter-module dependencies: 232 edges
circular dependencies: 4 circular dependency paths detected:
  1. detekt-api -> detekt-test -> detekt-api
  2. detekt-test -> detekt-test-utils -> detekt-test
  3. detekt-test -> detekt-test-utils -> detekt-test-junit -> detekt-test
  4. detekt-test-utils -> detekt-test-junit -> detekt-test-utils

top coupled modules (by afferent coupling Ca):
  • detekt-test: Ca=30, Ce=2, Instability I=0.06
  • detekt-api: Ca=29, Ce=3, Instability I=0.09
  • detekt-test-utils: Ca=25, Ce=5, Instability I=0.17
  • detekt-psi-utils: Ca=14, Ce=4, Instability I=0.22

The dependency scan immediately revealed 4 circular dependency paths. Unlike architectural entanglements in monolithic systems where production services cross-reference business domains, Detekt’s cycles are strictly isolated within test-harness infrastructure. Testing utilities (detekt-test, detekt-test-utils) require detekt-api to construct test findings, while test suites inside detekt-api simultaneously consume the test runners to validate API contracts.

Outside the test cluster, production modules maintain pristine architectural boundaries. Low-level AST parsing utilities in detekt-psi-utils (Cₐ=14, Cₑ=4) and foundational core types in detekt-utils (Cₐ=6, Cₑ=0, I=0.00) serve as unidirectional leaves, preventing compiler plugin internals from leaking into consumer rule sets.

Rule Specification Clones: Parameterized Linting Harnesses

Because Detekt implements dozens of independent lint rules—each accompanied by negative and positive test assertions—testing suites risk extensive copy-paste duplication. We executed prod-code duplicates to measure structural token replication across the codebase:

$ prod-code duplicates
scanned 1,175 files (135,761 lines of code)
found 20 clone groups across 112 files (9.4% codebase duplication)

top duplication clusters:
  • Clone Group #2627: 6 lines | 288 occurrences (Type-2 Parameterized)
    └─ rule specification assertions across detekt-rules-* test suites
       val actual = subject.lint(code)
       assertThat(actual).hasSize(1)
  • Clone Group #1011: 6 lines | 37 occurrences (Type-2 Parameterized)
    └─ multi-line string trimIndent and negative assertion verification
  • Clone Group #1534: 8 lines | 24 occurrences
    └─ Kotlin compiler environment setup and test fixture teardown

Detekt exhibits a 9.4% duplication rate, higher than typical runtime libraries. Crucially, examining individual clones demonstrates that this duplication is intentional: Clone Group #2627 accounts for 288 occurrences of parameterized test assertions across rule test suites.

Each test isolates a Kotlin code snippet in a raw string, passes it to subject.lint(code), and asserts findings count. Deduplicating this test scaffolding into shared helpers would reduce test readability. AST clone harvesting allows engineers to distinguish harmless test pattern replication from algorithmic drift in production visitors.

Structural Metavariables: Precondition Audits Across 31 Files

Static analyzers process untrusted user code and arbitrary syntax trees, making rigorous defensive validation essential. In Detekt, input arguments and configuration properties are validated using Kotlin’s require(condition) { "message" } contract.

Using prod-code structural-search, we queried the workspace for the AST pattern require($A) { $B }, where metavariables $A and $B bind to arbitrary boolean expressions and message closures:

$ prod-code structural-search 'require($A) { $B }'
scanned 1,179 files in 267.38 ms
found 58 matches across 31 files:
  • detekt-api/src/main/kotlin/dev/detekt/api/Finding.kt:16:9
    require(message.isNotBlank()) { "The message should not be empty" }
    └─ [$A = message.isNotBlank(), $B = "The message should not be empty"]
  • detekt-api/src/main/kotlin/dev/detekt/api/Location.kt:72:9
    require(line > 0) { "The source location line must be greater than 0" }
    └─ [$A = line > 0, $B = "The source location line must be greater than 0"]
  • detekt-api/src/main/kotlin/dev/detekt/api/internal/Signatures.kt:73:5
    require(startOffset < endOffset) { "Error building function signature..." }
    └─ [$A = startOffset < endOffset, $B = "Error building function signature..."]

To test automated transformation workflows across the codebase, we supplied this query to prod-code codemod to evaluate transforming argument preconditions into internal state checks:

$ prod-code codemod 'require($A) { $B } ==>> check($A) { $B }'
`require($A) { $B } ==>> check($A) { $B }`
154 changed line(s) in 31 file(s)

sample diff (Location.kt):
-        require(line > 0) { "The source location line must be greater than 0" }
-        require(column > 0) { "The source location column must be greater than 0" }
+        check(line > 0) { "The source location line must be greater than 0" }
+        check(column > 0) { "The source location column must be greater than 0" }

nothing was written; pass `apply: true` to make these edits

In 267 milliseconds, the structural search scanned over 1,100 files, identified every precondition site regardless of line breaks, and previewed a clean 154-line AST rewrite across 31 files without regex mismatches.

Polyglot Kotlin AST Refactoring and Cluster Language Server Verification

The cornerstone of prod-code’s architecture is executing surgical code transformations and validating them with a live compiler language server before changes reach the disk. In detekt-core/src/main/kotlin/dev/detekt/core/baseline/BaselineHandler.kt, the XML parser processes baseline suppression tags inside endElement:

override fun endElement(uri: String, localName: String, qName: String) {
    if (qName == ID) {
        check(content.isNotBlank()) { "The content of the ID element must not be empty" }
        when (current) {
            MANUALLY_SUPPRESSED_ISSUES -> manuallySuppressedIssues.add(content)
            CURRENT_ISSUES -> currentIssues.add(content)
        }
        content = ""
    }
}

We targeted the when (current) conditional branch on lines 26 to 29 and invoked prod-code extract-function to extract the issue registration logic:

$ prod-code extract-function \
    detekt-core/src/main/kotlin/dev/detekt/core/baseline/BaselineHandler.kt 26 13 \
    --to 29:14 \
    --name recordIssueByCurrentType

`fn recordIssueByCurrentType` extracted (detekt-core/.../BaselineHandler.kt);
the selection now reads `recordIssueByCurrentType(content)`

--- a/detekt-core/src/main/kotlin/dev/detekt/core/baseline/BaselineHandler.kt
+++ b/detekt-core/src/main/kotlin/dev/detekt/core/baseline/BaselineHandler.kt
@@ -25,6 +25,3 @@
             check(content.isNotBlank()) { "The content of the ID element must not be empty" }
-            when (current) {
-                MANUALLY_SUPPRESSED_ISSUES -> manuallySuppressedIssues.add(content)
-                CURRENT_ISSUES -> currentIssues.add(content)
-            }
+            recordIssueByCurrentType(content)
             content = ""
@@ -32,2 +29,9 @@
     }
+    private fun recordIssueByCurrentType(content: Any) {
+        return when (current) {
+            MANUALLY_SUPPRESSED_ISSUES -> manuallySuppressedIssues.add(content)
+            CURRENT_ISSUES -> currentIssues.add(content)
+        }
+    }

the analyzer accepts the result: 0 errors
nothing was written; pass `apply: true` to make this edit

The refactoring engine parsed the Kotlin AST, correctly deduced parameter usage (content: Any), preserved private class scoping, and placed the extracted method within BaselineHandler. The remote kotlin-language-server running on an isolated cluster node verified the modification in a shadow workspace and confirmed 0 compilation diagnostics. The entire analysis and verification loop executed remotely, delivering definitive compiler feedback to the local terminal in under two seconds.

A static analysis framework must hold its own codebase to the standards it enforces; verify dependency DAGs across test modules and let compiler language servers validate every AST transformation.

Cite this article
Citation
Alexander Panasenko (2026-09-30). Detekt Under the Microscope: What 67 Remote AST Tools Found Inside Kotlin's Static Code Analyzer. https://prod.codes/blog/detekt-under-the-microscope-67-ast-tools/