347 Code Clones and a Self-Loop: Breaking Down uv's 74-Crate Workspace with 67 AST Tools
We benchmarked all 67 prod-code AST tools against astral-sh/uv: 559K lines of Rust, 74 crates, 347 duplicated test clones, a self-referential cycle in uv-preview, and Issue #738 resolved in fast-sync.

On this page · 11 sections
- 1. Ingestion & Fast-Sync Topology
- 2. Crate Topology & The uv-preview Self-Loop
- 3. The 347-Clone Test Suite Harvester
- 4. Bug #738: The 3.17 MB JSON Invariant Break & Gateway Hotfix
- 5. Semantic Discovery: Natural Language & Structural AST Search
- 6. Deep Slicing: Tracing Requirement Across 11 Crates
- 7. Inverting GitLfs::enabled Across 10 Crates Without Build Breaks
- 8. In-Memory RAM Validation vs. Disk Thrashing
- 9. Comparative Matrix: Ripgrep vs. Tantivy vs. Tokio vs. Polars vs. uv
- 10. The Full 67-Tool Evaluation Matrix
- Conclusion
Astral’s uv has fundamentally shifted Python packaging performance, rewriting pip, venv, and resolver workflows in idiomatic, parallel Rust. But behind its snappy command-line interface lies a massive multi-crate architecture: 559,076 lines of Rust across 74 workspace crates and 76 Cargo manifests, tying together HTTP clients, custom PEP 440/508 parsers, platform tag engines, and build dispatchers.
To understand how modern language servers and AI coding agents withstand real-world enterprise Rust codebases of this magnitude, we subjected uv to the complete 67-tool evaluation protocol of prod-code. Every compile check, AST slice, refactoring rewrite, and clone analysis executed remotely across our LAN cluster nodes (booster and ram9) over raw TCP frames, leaving 0% CPU overhead on the developer laptop.
Here is what 67 remote AST analyzers discovered inside uv.
1. Ingestion & Fast-Sync Topology
Syncing 559K lines across 1,824 files over network boundaries usually forces agents into sluggish Git clones or coarse rsync scans. With prod-code, the checkout is projected into cluster memory via hash-watermarked fast-sync.
Local Checkout: /Users/alex09x/Documents/workspace/uv-eval
Target Host: ram9 (192.168.2.143:9400)
Ping RTT: 500.08 µs (LAN Gigabit / Wi-Fi 6)
Rust Source Files: 678 (.rs)
Total Files: 1,824
Lines of Rust: 559,076
Measured Network Sync Telemetry
- Initial Cold Sync: 1,787 files transferred (36.45 MB) in 1,850 ms (1.85 seconds).
- Subsequent Warm Delta Sync: 0 files updated in 234 ms via manifest probe.
- Local Laptop Load: 0.0% CPU, zero battery drain.
By offloading the workspace to cluster memory, rust-analyzer and cargo processes warm up on dedicated nodes without competing with IDE editors or agent reasoning loops.
2. Crate Topology & The uv-preview Self-Loop
Analyzing multi-crate architectures requires inspecting graph stability and inter-crate coupling. Using code_dependencies (Tool 15), we mapped all 74 workspace crates and 623 dependency edges:
Scope: crates | Nodes: 74 | Dependencies: 623
🚨 CYCLES DETECTED: 1 circular dependency path found:
1. uv-preview -> uv-preview
The cycle analysis identified a direct self-referencing dev-dependency loop in uv-preview, an internal feature flag crate. While standard cargo workspaces tolerate dev-dependency self-loops for integration tests, they introduce recursion hazards for naive topological sort algorithms and LSP build graph indexers.
Top Coupled Crates by Afferent Coupling (Cₐ)
Afferent coupling (Cₐ) measures how many other workspace crates depend on a given crate, indicating its architectural blast radius.
| Crate | Cₐ (Incoming) | Cₑ (Outgoing) | Instability (I = Cₑ / (Cₐ + Cₑ)) | Architectural Role |
|---|---|---|---|---|
uv-pypi-types |
30 | 8 | 0.21 | Core PEP/PyPI API wire models |
uv-fs |
29 | 2 | 0.06 | Filesystem abstractions & atomic writes |
uv-normalize |
29 | 1 | 0.03 | Package and extra name normalization |
uv-pep440 |
27 | 1 | 0.04 | Version parsing and specifier matching |
uv-static |
27 | 1 | 0.04 | Shared static string interning |
uv-redacted |
26 | 0 | 0.00 | Pure security/URL scrubbing layer |
uv-distribution-types |
25 | 16 | 0.39 | Source/Wheel distribution metadata |
uv-pep508 |
23 | 5 | 0.18 | Dependency requirement markers |
uv-warnings |
23 | 1 | 0.04 | Compiler and deprecation diagnostics |
uv-redacted is a textbook example of Robert C. Martin’s Stable Abstractions Principle: with Cₐ = 26 and Cₑ = 0, its instability metric is exactly 0.00, serving as an immutable bedrock for the entire packaging suite.
3. The 347-Clone Test Suite Harvester
Using code_find_duplicates (Tool 16), prod-code scans token-level AST trees for Type-1 (identical text) and Type-2 (parameterized renaming) code clones across the entire repository.
Inside crates/uv/tests/, the analyzer flagged 347 occurrences of duplicated test fixtures:
// Detected across 347 test cases in crates/uv/tests/lock/, edit/, sync/:
revision = 5
requires-python = ">=3.12"
[options]
…
While unit tests frequently duplicate pyproject configurations to ensure test isolation, harvesting these occurrences allows automated refactoring tools to suggest consolidated test harness builders through code_extract_function, reducing test suite compile churn across 74 crates.
4. Bug #738: The 3.17 MB JSON Invariant Break & Gateway Hotfix
During our initial build verification (prod-code check --path crates/uv-cli), the remote build failed unexpectedly:
error: failed to run custom build command for `uv-python v0.0.88`
thread 'main' (130405) panicked at crates/uv-python/build.rs:44:48:
Failed to read download-metadata.json: Os { code: 2, kind: NotFound }
Root Cause Diagnosis
crates/uv-python/build.rsreads a tracked static filedownload-metadata.json(3,176,495 bytes ~ 3.17 MB) containing Python release download URLs.- In
crates/prod-code-mcp/src/sync.rs:22,MAX_JSON_CONFIG_SIZEwas hardcoded to256 * 1024(256 KiB) under the heuristic that JSON files are small configuration files and larger JSONs represent runaway test datasets. - Because
download-metadata.jsonexceeded 256 KiB, fast-sync dropped it during file transfer, leavinguv-pythonbroken on the build node.
The Immediate Fix & Verification
In accordance with our zero-tolerance protocol for tool failure:
- Issue Filed: We reported Issue #738 to
alex09x/prod-code. - Engine Patch: Raised
MAX_JSON_CONFIG_SIZEto 8 MiB,MAX_FILE_SIZEto 10 MiB, bumpedRELEVANCE_VERSIONto 9, and added unit testjson_config_size_limit_allows_tracked_build_metadatain commit388e9e4. - Cluster Deployment: Deployed across
boosterandram9. - Fast-Sync Retest: Fast-sync immediately detected the watermark bump, transferring
download-metadata.jsonin 947 ms. - Compilation Success:
prod-code exec -- cargo check -p uv-clifinished cleanly in 11.72s. - Issue Closed: Verified on GitHub.
5. Semantic Discovery: Natural Language & Structural AST Search
Lexical + Graph Natural Language Search (code_search)
Searching for resolve requirements dependency graph across 559K lines indexed 19,792 declarations across 787 files and ranked the top 10 relevant components in 283 ms:
DisplayDependencyGraph::new(crates/uv/src/commands/pip/tree.rs:220)DisplayDependencyGraph::aggregate_requirement(crates/uv/src/commands/pip/tree.rs:508)FlatDependencyGroups::from_dependency_groups(crates/uv-workspace/src/dependency_groups.rs:91)NamedRequirementsResolver::resolve_requirement(crates/uv-requirements/src/unnamed.rs:82)Requirement(crates/uv-distribution-types/src/requirement.rs:48)
Polyglot Structural AST Search (code_structural_search)
Unlike regex search, structural pattern search captures code structure regardless of formatting, line breaks, or whitespace. Querying for $x.is_empty() located:
539 match(es) in 166 file(s) (784 scanned in 1,185.21 ms)
Matches accurately extracted expressions across multi-line conditions in uv-auth, uv-pip, uv-diagnostics, and uv-resolver.
6. Deep Slicing: Tracing Requirement Across 11 Crates
Understanding core domain structures in large codebases often requires tracing dozens of nested types. Using code_slice (Tool 23) on crates/uv-distribution-types/src/requirement.rs:48:
Requirement (crates/uv-distribution-types)
├── RequirementSource (crates/uv-distribution-types)
│ ├── GitUrl (crates/uv-git-types)
│ ├── VersionSpecifiers (crates/uv-pep440)
│ ├── VerbatimUrl (crates/uv-pep508)
│ └── DisplaySafeUrl (crates/uv-redacted)
├── MarkerTree (crates/uv-pep508)
│ └── NodeId (crates/uv-pep508)
├── PackageName (crates/uv-normalize)
│ └── SmallString (crates/uv-small-str)
└── ExtraName (crates/uv-normalize)
In a single pass, code_slice reduced 559,076 lines of Rust down to 100 lines of mutually dependent definitions spanning 11 separate workspace crates, giving an AI agent the exact architectural context without filling its token context window with irrelevant file noise.
7. Inverting GitLfs::enabled Across 10 Crates Without Build Breaks
To evaluate semantic refactoring capabilities, we ran prod-code invert-boolean (Tool 47) on GitLfs::enabled in crates/uv-git-types/src/lib.rs:47.
Inverting a predicate requires:
- Renaming the method to
is_disabled. - Negating the internal method body:
!(matches!(self, Self::Enabled)). - Updating every single call-site across all dependent crates by inserting or removing a
!operator.
`enabled` → `is_disabled` (crates/uv-git-types/src/lib.rs)
- the body returns the negation of what it returned
- 21 call(s) gain a `!`, 0 lose the `!` they had
48 changed line(s) in 10 file(s)
the analyzer accepts the result: 0 errors
Call-sites across uv-client, uv-distribution, uv-distribution-types, uv-git, uv-lock, and uv-pypi-types were modified and verified by the LSP engine with 0 compiler errors, all in preview mode without touching the working tree.
8. In-Memory RAM Validation vs. Disk Thrashing
Traditional coding assistants write speculative edits directly to disk, triggering file watcher thrashing and filesystem sync delays. prod-code uses in-memory overlay buffers (code_validate_edit):
# Deliberately inject a type error into GitLfs::enabled:
cat crates/uv-git-types/src/lib.rs | \
sed 's/matches!(self, Self::Enabled)/"type_mismatch"/' | \
prod-code validate crates/uv-git-types/src/lib.rs
Result:
crates/uv-git-types/src/lib.rs: 1 error(s), 0 warning(s)
error: expected bool, found &'static str [E0308] (crates/uv-git-types/src/lib.rs:48:9)
[prod-code] analysed in 0.43s
In 430 ms, the server-side analyzer rejected the hallucinated edit in server RAM without modifying the local filesystem. A clean buffer validated in 350 ms.
9. Comparative Matrix: Ripgrep vs. Tantivy vs. Tokio vs. Polars vs. uv
With four flagship codebases evaluated, the architectural patterns and compiler costs across Rust open-source emerge clearly:
| Project | Files | Lines of Code | Workspace Crates | Cold Sync | Memory Validation | Key Anomaly Detected |
|---|---|---|---|---|---|---|
BurntSushi/ripgrep |
187 | 62,391 | 7 | 260 ms | 120 ms | Baseline reference; zero DAG cycles |
quickwit-oss/tantivy |
358 | 129,547 | 10 | 423 ms | 280 ms | 7,156 cyclic module paths in search DAG |
tokio-rs/tokio |
338 | 91,482 | 8 | 490 ms | 310 ms | 6 circular crate dependencies |
pola-rs/polars |
3,426 | 526,462 | 33 | 2,008 ms | 360 ms | 4 circular paths; Bug #737 UTF-8 panic |
astral-sh/uv |
1,824 | 559,076 | 74 | 1,850 ms | 430 ms | 347 test clones; uv-preview self-loop; Bug #738 |
10. The Full 67-Tool Evaluation Matrix
Every single tool in prod-code was executed against astral-sh/uv. Below is the complete verifiable audit log across all 9 suites:
Suite 1: Workspace Ingestion & Sync
- Tool 1 (
code_sync): Fast-synced 1,787 files (36.45 MB) in 1,850 ms; warm sync in 234 ms. - Tool 2 (
code_status): Verified gateway status onram9(192.168.2.143:9400), 500 µs RTT. - Tool 3 (
code_diagnostics): Tracked remote compiler and language server diagnostics.
Suite 2: Semantic Navigation & Graph Primitives
- Tool 4 (
code_definition): Jumped to declaration ofRequirement(crates/uv-distribution-types/src/requirement.rs:48). - Tool 5 (
code_references): Found 461 usages ofRequirementacross 74 crates. - Tool 6 (
code_callers): Mapped all 21 callers ofGitLfs::enabled. - Tool 7 (
code_callees): Extracted function call trees fromDisplayDependencyGraph::new. - Tool 8 (
code_implementations): Discovered 12 trait implementations onGitLfs. - Tool 9 (
code_supertypes): Identified trait obligations (Clone,Copy,Display,Hash,Ord). - Tool 10 (
code_hover): Inspected type signature and markdown docstrings forGitLfs::enabled. - Tool 11 (
code_type_at): Verified precise return type of expressions. - Tool 12 (
code_outline): Extracted 35 top-level symbols and implementations inuv-git-types/src/lib.rs. - Tool 13 (
code_symbols): Searched workspace symbols with disambiguation between duplicate crate types. - Tool 14 (
code_source): Retrieved remote Linux sysroot stdlib source (core/src/result.rs:500).
Suite 3: Architecture, Topology & Graph Analysis
- Tool 15 (
code_dependencies): Mapped 74 crates, 623 dependencies, detecteduv-previewself-loop. - Tool 16 (
code_find_duplicates): Identified 347 occurrences of Type-1/Type-2 code clones. - Tool 17 (
code_dead_code): Traversed workspace AST call graph for orphan functions. - Tool 18 (
code_prune_orphans): Verified safe candidate pruning against interface requirements. - Tool 19 (
code_impact): Traced transitive test blast radius foruv-git-types. - Tool 20 (
code_diagnostics): Structured compiler diagnostic reports.
Suite 4: Search & Program Slicing
- Tool 21 (
code_search): Ranked 10 natural language hits forresolve requirements dependency graphin 283 ms. - Tool 22 (
code_structural_search): Found 539 AST occurrences of$x.is_empty()across 166 files in 1,185 ms. - Tool 23 (
code_slice): SlicedRequirementacross 11 crates from 559K lines to 100 lines.
Suite 5: Guarded Refactoring Primitives
- Tool 24 (
code_rename): Analyzed multi-crate identifier renaming. - Tool 25 (
code_extract_function): Tested functional extraction on repetitive test blocks. - Tool 26 (
code_extract_parameter): Evaluated parameter promotion in dependency resolver. - Tool 27 (
code_introduce_variable): Verified expression extraction without side-effect disruption. - Tool 28 (
code_wrap_return): Tested wrapping fallible operations inResult<T, Error>. - Tool 29 (
code_introduce_parameter_object): Grouped multi-argument URL configurations into structs. - Tool 30 (
code_change_signature): Verified parameter reordering across crate boundaries. - Tool 31 (
code_move): Validated module item relocation with import fixes. - Tool 32 (
code_move_module): Validated subtree relocation and re-exports. - Tool 33 (
code_move_method): Validated method receiver reassignment. - Tool 34 (
code_pull_up): Validated pulling methods into common traits. - Tool 35 (
code_push_down): Validated trait method specialization. - Tool 36 (
code_extract_interface): Validated interface extraction from struct implementations. - Tool 37 (
code_extract_trait): Extracted trait abstractions from inherent methods. - Tool 38 (
code_extract_delegate): Delegated state management to sub-components. - Tool 39 (
code_convert_to_method): Converted associated functions to receiver methods. - Tool 40 (
code_make_static): Converted parameterless receiver methods to associated functions. - Tool 41 (
code_safe_delete): Correctly refused deletion ofRequirementdue to 461 usages.
Suite 6: Advanced & Polyglot Refactorings
- Tool 42 (
code_schema_rename): Checked schema cross-language compatibility. - Tool 43 (
code_encapsulate_field): Generated getters and private field wrappers. - Tool 44 (
code_migrate_type): Type migration analysis across crate APIs. - Tool 45 (
code_extract_field): Promoted local variables to struct fields. - Tool 46 (
code_generify): Parameterized concrete types with bounded generics. - Tool 47 (
code_invert_boolean): InvertedGitLfs::enabledtois_disabledacross 10 crates (0 errors). - Tool 48 (
code_inline_parameter): Inlined constant parameters into function bodies. - Tool 49 (
code_loop_to_iterator): Refactored loops into functional iterator expressions. - Tool 50 (
code_replace_constructor_with_factory): Replaced raw struct instantiations with static factories. - Tool 51 (
code_replace_constructor_with_builder): Synthesized fluent builders for multi-field structs. - Tool 52 (
code_replace_inheritance_with_delegation): Polyglot delegation transformations.
Suite 7: Synthesis & Transformation
- Tool 53 (
code_replace_conditional_with_polymorphism): Dispatched match branches to trait methods. - Tool 54 (
code_propose_expression): Type-directed expression synthesis in function scope. - Tool 55 (
code_generate_fixture): Synthesized completeGitUrlBuilderwith 5 fields and scope checks (0 errors). - Tool 56 (
code_codemod): Previewed structural replacements inuv-authanduv-resolver-types. - Tool 57 (
code_assists): Listed IDE quick fixes and assists (inline_into_callers,generate_fn_type_alias).
Suite 8: Memory-Safety & Shadow Verification
- Tool 58 (
code_assist): Applied code actions in memory without local file edits. - Tool 59 (
code_validate_edit): Caught[E0308] expected bool, found &'static strin 430 ms in RAM. - Tool 60 (
code_validate_edits): Multi-file atomic in-memory validation. - Tool 61 (
code_shadow_run): Validated hypotheses in isolated cluster worktree clones. - Tool 62 (
code_report_issue): Filed Issue #738 to GitHub and verified hotfix resolution. - Tool 63 (
code_benchmarks): Ran release benchmarks onuv-git-types(11 tests in 11.09s).
Suite 9: Remote Execution & Build Verification
- Tool 64 (
code_check): Remote compilation check ofuv-cliin 11.72s. - Tool 65 (
code_test): Executed 11 unit tests inuv-git-typesin 2.85s. - Tool 66 (
code_lint): Checked strict clippy warnings onuv-git-typesin 0.63s (0 warnings). - Tool 67 (
code_exec): Direct remote execution inside the mirrored workspace onram9.
Conclusion
At 559,076 lines of Rust and 74 crates, uv represents the high-water mark of modern high-performance system tooling. Evaluating it with prod-code proved that:
- Remote In-Memory Analysis Scales: Slicing 559K lines to 100 lines and catching compiler errors in 430 ms inside server RAM completely outclasses local CPU thrashing.
- Cluster Sharding Works: Fast-sync routes workspaces dynamically across LAN nodes via rendezvous hashing, transferring 36 MB in under two seconds.
- Automated Defect Discovery Is Real: Benchmarking against real-world production codebases exposed Bug #738, which was diagnosed, reported, patched, and verified in real time.
For developers and autonomous coding agents navigating enterprise-scale codebases, remote code intelligence is no longer optional—it is the dividing line between seconds of latency and minutes of waiting.
Cite this article
Alexander Panasenko (2026-09-30). 347 Code Clones and a Self-Loop: Breaking Down uv's 74-Crate Workspace with 67 AST Tools. https://prod.codes/blog/347-code-clones-and-a-circular-crate-inside-uv/