notes · · 3 min

Guava Under the Microscope: What Remote AST Tools Found Inside Google's Java Core

We benchmarked prod-code AST tools against google/guava: 795,118 lines of Java, 5-module DAG resolution in 100ms, and 33-occurrence clone analysis.

On this page · 4 sections
  1. The 795,118-Line Monolith and its 5-Module Dependency DAG
  2. Token-Level Duplication: The 33-Occurrence Clone Phenomenon
  3. Metavariable Codemods Across 3,279 Files in 545.82ms
  4. The Architectural Takeaway: Java Monorepos Need AST-Aware Toolchains

Google’s Guava is one of the most widely deployed utility libraries in the Java ecosystem. From immutable collections and concurrency primitives to cache interfaces and hashing algorithms, Guava powers enterprise backends across hundreds of thousands of organizations. But beneath its familiar ImmutableList and Preconditions surfaces lies a massive codebase developed over nearly two decades.

Analyzing enterprise Java at this scale tests code intelligence tooling against deep module structures, heavy parameter overloading, and sprawling package boundaries. The repository contains 795,118 lines of Java across 3,275 source files. We executed prod-code’s remote AST toolchain against Guava HEAD, validating multi-module DAG extraction, token-level duplication harvesting, metavariable structural search, and workspace-wide AST codemods across distributed cluster nodes.

The 795,118-Line Monolith and its 5-Module Dependency DAG

Large Java repositories often struggle with architectural drift: circular package dependencies, sprawling module boundaries, and hidden cross-module coupling. In Maven and Gradle projects, understanding the real coupling between modules requires analyzing both build descriptors and actual source-level import structures.

We executed prod-code dependencies against Guava. The analysis traversed the multi-module Maven tree, mapped inter-module references across all 3,275 source files, and synthesized the module dependency graph in 100 milliseconds:

$ prod-code dependencies
prod-code Architecture & Dependency Graph Analysis: /tmp/guava-eval
────────────────────────────────────────────────────────────────────────────────
Found 5 modules across 3,275 source files (analyzed in 100ms)

Module Coupling (Afferent/Efferent):
  • guava: Ca=4, Ce=0, instability=0.00
  • guava-testlib: Ca=3, Ce=1, instability=0.25
  • guava-gwt: Ca=1, Ce=3, instability=0.75
  • guava-tests: Ca=1, Ce=2, instability=0.67
  • guava-bom: Ca=0, Ce=3, instability=1.00

Dependency Cycles:
  No package or module dependency cycles detected! Clean DAG architecture.

The output confirms a disciplined unidirectional dependency architecture:

  1. guava sits at the foundational layer with an Afferent Coupling (Ca) of 4 and Efferent Coupling (Ce) of 0, yielding an instability score of 0.00. Nothing beneath it is imported, and all other modules depend directly upon it.
  2. guava-testlib acts as the shared testing harness with Ca=3 and Ce=1, depended on by guava-gwt and guava-tests.
  3. guava-bom serves as the root Bill of Materials with Ce=3 and Ca=0, maintaining pure outward packaging metadata.
  4. Across the entire 795,118-line monorepo, zero architectural cycles exist between modules.

Token-Level Duplication: The 33-Occurrence Clone Phenomenon

In high-performance Java libraries, generic collections frequently avoid heap allocation for varargs arrays by declaring discrete overloads for small cardinalities: of(e1), of(e1, e2), up to of(e1, e2, e3, e4, e5, e6, e7, e8, e9, e10). While ergonomically convenient for callers, this pattern produces significant token duplication that traditional copy-paste detectors struggle to categorize cleanly.

We executed prod-code duplicates with a 20-token window across all 795,118 lines. The engine harvested 20 major clone groups:

$ prod-code duplicates
prod-code Duplicate Code Analysis (Token-level sliding window)
─────────────────────────────────────────────────────────────────
Found 20 clone groups across 795,118 lines (scanned 3,275 files)

Clone Group #222561: 33 occurrences (~3 lines each, 21 tokens)
  • guava-gwt/src-super/com/google/common/collect/super/.../RegularImmutableMap.java:36
  • guava-gwt/src-super/com/google/common/collect/super/.../RegularImmutableMap.java:42
  • guava-gwt/src-super/com/google/common/collect/super/.../RegularImmutableMap.java:48
  Pattern: // To avoid J2ObjC optimizer bug, create the instance directly

Clone Group #215553: 33 occurrences (~10 lines each, 68 tokens)
  • guava/src/com/google/common/collect/ImmutableMap.java:124
  • guava/src/com/google/common/collect/ImmutableBiMap.java:98
  • guava/src/com/google/common/collect/ImmutableSortedMap.java:152
  Pattern: ImmutableMap.of(k1, v1, k2, v2, ...) parameter validation overload chain

The clone detector surfaced two distinct structural engineering realities in Guava:

  • Clone Group #222561 reveals 33 identical constructor wrappers across GWT/J2ObjC emulation modules, explicitly bypassing compiler optimization traps on transpilations.
  • Clone Group #215553 captures the systematic overload hierarchy across ImmutableMap, ImmutableBiMap, and ImmutableSortedMap, ensuring type safety across arbitrary fixed-arity maps without boxing penalties.

Metavariable Codemods Across 3,279 Files in 545.82ms

Refactoring modern Java codebases often entails migrating legacy assertions or utility methods to modern Java 9+ standard library equivalents. For instance, replacing Preconditions.checkNotNull(x) with Objects.requireNonNull(x) improves JIT inlining characteristics and reduces third-party library dependencies.

Using string search or regular expressions for such rewrites risks invalid transformations on multi-line expressions or nested argument boundaries. A structural AST engine matches syntax trees directly, capturing metavariable expressions regardless of formatting or whitespace.

We executed prod-code structural-search to find all explicit Preconditions.checkNotNull($A) call sites across the codebase:

$ prod-code structural-search --path . 'Preconditions.checkNotNull($A)'
prod-code Structural AST Search: `Preconditions.checkNotNull($A)`
────────────────────────────────────────────────────
92 match(es) in 44 file(s) (3,279 scanned in 545.82ms)

  • android/guava/src/com/google/common/base/JdkPattern.java:29:20
    Preconditions.checkNotNull(pattern)
    └─ [$A = pattern]
  • android/guava/src/com/google/common/base/Platform.java:91:5
    Preconditions.checkNotNull(pattern)
    └─ [$A = pattern]
  • android/guava/src/com/google/common/collect/Queues.java:325:5
    Preconditions.checkNotNull(buffer)
    └─ [$A = buffer]
  • android/guava/src/com/google/common/util/concurrent/Futures.java:1101:5
    Preconditions.checkNotNull(callback)
    └─ [$A = callback]

In 545.82 milliseconds, the engine traversed 3,279 files and isolated 92 invocations across 44 files, extracting parameter expressions into metavariable $A.

Next, we evaluated the structural codemod transformation Preconditions.checkNotNull($A) ==>> Objects.requireNonNull($A) in dry-run mode:

$ prod-code codemod 'Preconditions.checkNotNull($A) ==>> Objects.requireNonNull($A)' --path .
`Preconditions.checkNotNull($A) ==>> Objects.requireNonNull($A)`
184 changed line(s) in 44 file(s)

--- a/android/guava/src/com/google/common/base/JdkPattern.java
+++ b/android/guava/src/com/google/common/base/JdkPattern.java
@@ -27,5 +27,5 @@
   JdkPattern(Pattern pattern) {
-    this.pattern = Preconditions.checkNotNull(pattern);
+    this.pattern = Objects.requireNonNull(pattern);
   }

--- a/android/guava/src/com/google/common/collect/Queues.java
+++ b/android/guava/src/com/google/common/collect/Queues.java
@@ -323,5 +323,5 @@
     TimeUnit unit)
     throws InterruptedException {
-    Preconditions.checkNotNull(buffer);
+    Objects.requireNonNull(buffer);

The codemod generated 184 changed lines across 44 files, preserving exact code indentations, comment alignments, and AST fidelity without writing corrupting diffs or misinterpreting overloaded two-argument check variants (checkNotNull(x, message)).

The Architectural Takeaway: Java Monorepos Need AST-Aware Toolchains

Evaluating 795,118 lines of Google Guava highlights the boundary where text-based developer tooling breaks down in mature enterprise repositories:

  1. Topological clarity beats raw line metrics: A 795,118-line monorepo remains maintainable when module boundaries form a strict unidirectional acyclic graph. Guava’s clean instability metric (0.00 at root) demonstrates that foundational libraries must never leak inward dependencies.
  2. Structural pattern engines eliminate rewrite hazards: Migrating enterprise Java APIs cannot rely on regex search-and-replace. Token-aware AST pattern matching isolates semantic arguments in half a second across thousands of files, generating clean, reviewable diffs.
  3. Cluster offloading maintains developer velocity: Indexing 3,275 files and computing graph centralities consumes substantial memory and parsing bandwidth. Offloading language intelligence, clone detection, and verification to remote cluster nodes keeps developer workstations cool and responsive.

Architectural Rule: In mature enterprise codebases, enforce strict afferent/efferent layer hierarchies where foundational modules maintain zero efferent coupling, and execute API migrations exclusively via AST-level metavariable transformations rather than lexical pattern substitutions.

Cite this article
Citation
Alexander Panasenko (2026-09-30). Guava Under the Microscope: What Remote AST Tools Found Inside Google's Java Core. https://prod.codes/blog/guava-under-the-microscope-67-ast-tools/