Black Under the Microscope: What 67 AST Tools Found Inside Python's Uncompromising Formatter
We benchmarked all 67 prod-code AST tools against psf/black: 135K lines of Python, 12-callsite parameter object bundling across 4 files, import cycle detection in concurrency, 99.8% slicing reduction, and mathematical safety proofs in predicate inversion.

On this page · 8 sections
- 1. Import Cycles and Architecture Graph
- 2. Type-2 Clone Clusters Across Python Feature Sets
- 3. Structural Pattern Matching in 119 Milliseconds
- 4. 3-Way Semantic Search Across 3,133 Declarations
- 5. Transitive AST Slicing: 99.8% Reduction
- 6. Automated Python Refactoring & Safety Proofs
- 7. Pre-Flight In-Memory Shadow Validation (prod-code validate)
- Summary Matrix
psf/black reshaped how the Python community writes software. By enforcing a deterministic, uncompromising code style and removing all manual debates over formatting, Black became the standard code formatter across millions of Python repositories worldwide.
Under the hood, Black is not a simple regex or token replacer. It is a full-blown compiler pipeline: it tokenizes source text via tokenize, parses it into concrete syntax trees using an internal fork of blib2to3, normalizes brackets and line lengths, splits lines according to delimiter priorities, and runs AST equivalence checks (assert_equivalent) against Python’s built-in ast module to guarantee that formatting never alters program semantics.
We evaluated all 67 prod-code AST tools against psf/black (commit a7c2368): 135,022 lines of Python across 358 files (18,716 lines in the core src/black/ package). The entire suite was executed against cluster build nodes (booster 192.168.2.168:9400 and ram9 192.168.2.143:9400), incurring 0% local laptop CPU.
Here is what the analysis revealed.
1. Import Cycles and Architecture Graph
Unlike compiled languages with strict module DAG enforcement, Python allows cyclical imports as long as symbols are resolved after module initialization or imported lazily inside functions.
Running prod-code diagnostics with basedpyright-langserver on src/black/__init__.py surfaced an immediate architectural finding:
error: Cycle detected in import chain [reportImportCycles] (src/black/__init__.py:1:1)
[prod-code] analysed in 1.45s
Tracing the cycle through prod-code refs and module traversal reveals how Black’s concurrency subsystem interacts with the root package:
flowchart LR
Init["src/black/__init__.py"] -->|"line 768: from black.concurrency import reformat_many"| Conc["src/black/concurrency.py"]
Conc -->|"line 24: from black import WriteBack, format_file_in_place"| Init
src/black/concurrency.py imports WriteBack and format_file_in_place directly from top-level black at module load time. Meanwhile, src/black/__init__.py imports reformat_many inside main() at line 768. Because the import in __init__.py is deferred to runtime execution, Python executes without an ImportError, but the circular dependency loop remains embedded in the module graph.
2. Type-2 Clone Clusters Across Python Feature Sets
Black targets every supported Python minor version from Python 3.9 through Python 3.14. In src/black/mode.py, feature flags describe language constructs supported by each target version.
Running prod-code duplicates --min-lines 10 revealed repetitive clone clusters across version definitions:
[Clone Group #432] 10 lines | 7 occurrences (Type-2 (Parameterized))
• Occurrence 1: src/black/mode.py:108-117
• Occurrence 2: src/black/mode.py:123-132
• Occurrence 3: src/black/mode.py:139-148
• Occurrence 4: src/black/mode.py:157-166
• Occurrence 5: src/black/mode.py:176-185
• Occurrence 6: src/black/mode.py:196-205
• Occurrence 7: src/black/mode.py:218-227
Preview:
│ Feature.NUMERIC_UNDERSCORES,
│ Feature.TRAILING_COMMA_IN_CALL,
│ Feature.TRAILING_COMMA_IN_DEF,
│ Feature.ASYNC_KEYWORDS,
Seven different version maps repeat identical subsets of Feature enum values, representing candidates for shared feature sets (PYTHON_36_PLUS_FEATURES, etc.).
3. Structural Pattern Matching in 119 Milliseconds
Searching Python code for structural idioms using regex is notoriously brittle due to arbitrary whitespace, multiline expressions, and optional parentheses.
We ran prod-code structural-search across src/black searching for all guard-and-raise patterns:
prod-code struct-search "if \$cond: raise \$exc" --path src/black
In 119.90 milliseconds, the engine parsed and evaluated all 25 files in src/black/, finding 35 matches in 8 files:
⚡ prod-code Structural AST Search: `if $cond: raise $exc`
────────────────────────────────────────────────────
35 match(es) in 8 file(s) (25 scanned in 119.90ms)
• src/black/__init__.py:1184:5 if src_contents == dst_contents:
└─ [$cond = src_contents == dst_contents, $exc = NothingChanged]
• src/black/__init__.py:1227:5 if dst == src:
└─ [$cond = dst == src, $exc = NothingChanged]
• src/black/linegen.py:976:5 if not matching_bracket or not tail_leaves:
└─ [$cond = not matching_bracket or not tail_leaves, $exc = CannotSplit]
• src/black/linegen.py:1068:5 if not (opening_bracket and closing_bracket and head_leaves):
└─ [$cond = not (opening_bracket and closing_bracket and head_leaves), $exc = CannotSplit]
The AST matcher extracted structural bindings directly: capturing when Black short-circuits with NothingChanged or halts line splitting with CannotSplit.
4. 3-Way Semantic Search Across 3,133 Declarations
Black’s line-splitting algorithm relies heavily on bracket tracking and delimiter priorities. To locate where delimiter priorities dictate line splits without knowing function names in advance, we ran prod-code search:
prod-code search "split line by delimiter priority"
In 8 milliseconds, across 3,133 declarations and 358 files, the 3-way RRF ranker (combining dense embeddings, BM25 lexical search, and typed call graph centrality) returned the exact core implementations:
10 hit(s) for `split line by delimiter priority` in 8 ms
1. [function] is_split_after_delimiter src/black/brackets.py:219 (cosine 0.844)
2. [function] is_split_before_delimiter src/black/brackets.py:244 (cosine 0.837)
3. [function] BracketTracker::delimiter_count_with_priority src/black/brackets.py:150 (cosine 0.794)
4. [function] BracketTracker::max_delimiter_priority src/black/brackets.py:142 (cosine 0.810)
5. [function] max_delimiter_priority_in_atom src/black/brackets.py:340 (cosine 0.771)
6. [function] should_split_line src/black/linegen.py:2279 (cosine 0.731)
7. [function] delimiter_split src/black/linegen.py:1505 (cosine 0.780)
Both dense semantic similarity and lexical keyword matching aligned on is_split_after_delimiter and delimiter_count_with_priority.
5. Transitive AST Slicing: 99.8% Reduction
When debugging why a line was or was not formatted, inspecting the entire 135K-line repository or even the 2,463-line linegen.py file is overwhelming.
We ran prod-code slice on the heuristic function can_be_split at src/black/lines.py:1483:
prod-code slice src/black/lines.py --line 1483 --depth 2
The slicer followed backward data-flow and type dependencies:
- Seed function
can_be_split(lines 1483–1516) Lineclass definition and its leaf accessors (src/black/lines.py)Modedataclass definition (src/black/mode.py)- Constant AST sets
OPENING_BRACKETSandCLOSING_BRACKETS(src/black/nodes.py) - Token constants
NAME,STRING,DOT(blib2to3/pgen2/token.py)
The result was an isolated, self-contained 240-line snippet containing everything can_be_split depends on and nothing else—reducing the cognitive surface by 99.8%.
6. Automated Python Refactoring & Safety Proofs
Predicate Inversion with Non-Call Reference Refusal
We tested inverting can_be_split into cannot_be_split:
prod-code invert-boolean src/black/lines.py --line 1483 --character 5 --to cannot_be_split
The engine executed a complete semantic analysis:
- Inverted 6 return paths in
can_be_split: five instances ofreturn Falseinverted toreturn True, and the trailingreturn Trueinverted toreturn False. - Rewrote call sites:
The leading# Before elif not can_be_split(rhs.body) and ...: # After elif cannot_be_split(rhs.body) and ...:notwas cleanly cancelled out. - Safety Refusal: The engine discovered a non-call reference:
Becausenot rewritten (1 reference(s) that are not a call — a function used as a value keeps its old meaning under its new name; nothing is written while any remains): src/black/linegen.py:32:5 `can_be_split` used as a valuecan_be_splitwas imported by name viafrom black.lines import can_be_splitatlinegen.py:32, the engine mathematically refused to commit the change until the import symbol itself was updated, preventing broken runtime imports.
Parameter Object Bundling Across 4 Files
In src/black/__init__.py:1724, the core equivalence assertion function takes two separate parameters:
def assert_equivalent(src: str, dst: str) -> None:
...
We tested bundling (src: str, dst: str) into a dataclass CodePair:
prod-code parameter-object --path src/black/__init__.py \
--param src --param dst --name CodePair assert_equivalent
The refactoring engine automatically:
- Synthesized an idiomatic
@dataclass:@dataclass class CodePair: """The parameters `assert_equivalent` takes together.""" src: str dst: str - Rewrote
assert_equivalent(code_pair: CodePair)to accesscode_pair.srcandcode_pair.dst. - Injected
from dataclasses import dataclassintosrc/black/__init__.py. - Rewrote 12 call sites across 4 files:
scripts/fuzz.py:black.assert_equivalent(black.CodePair(src=src_contents, dst=dst_contents))src/black/__init__.py:assert_equivalent(CodePair(src=src_contents, dst=dst_contents))tests/test_black.py:black.assert_equivalent(black.CodePair(src=source, dst=result.stdout))tests/util.py:black.assert_equivalent(black.CodePair(src=source, dst=actual))
Both local calls and module-qualified calls (black.assert_equivalent(...)) were correctly rewritten with module prefixes (black.CodePair(...)).
Function Extraction
We tested extracting the token check leaves[0].type == token.STRING and leaves[1].type == token.DOT at src/black/lines.py:1494 into a helper function is_string_dot_prefix:
prod-code extract-function src/black/lines.py 1494 8 --to 1494:70 \
--name is_string_dot_prefix
Output:
def is_string_dot_prefix(leaves):
return leaves[0].type == token.STRING and leaves[1].type == token.DOT
def can_be_split(line: Line) -> bool:
...
if is_string_dot_prefix(leaves):
call_count = 0
Verified against basedpyright: 0 errors.
7. Pre-Flight In-Memory Shadow Validation (prod-code validate)
Applying experimental edits directly to disk risks corrupting active developer checkouts. prod-code validate verifies proposed changes in memory on the remote node before anything touches disk.
We simulated an injected attribute typo in src/black/lines.py:
leaves = line.non_existent_attribute
Piping the modified file into prod-code validate:
cat src/black/lines.py | sed 's/leaves = line.leaves/leaves = line.non_existent_attribute/' | \
prod-code validate src/black/lines.py
Result:
src/black/lines.py: 1 error(s), 14 warning(s)
(8 diagnostic(s) the file already had before this edit are not counted)
error: Cannot access attribute "non_existent_attribute" for class "Line" [reportAttributeAccessIssue] (src/black/lines.py:1490:19)
[prod-code] analysed in 0.68s
The validator automatically subtracted the 8 pre-existing baseline diagnostics, flagged the single newly introduced attribute error, and completed analysis in 0.68 seconds in remote RAM.
Summary Matrix
| Metric | Result |
|---|---|
| Target Repository | psf/black (a7c2368) |
| Codebase Size | 135,022 lines of Python across 358 files |
| Cluster Node | booster (32 cores, 0.93 ms RTT) & ram9 (32 cores, 1.12 ms RTT) |
| Laptop CPU Usage | 0.0% |
| Import Cycle Analysis | Cycle detected in src/black/__init__.py via concurrency.py |
| Clone Clusters | 10-line Type-2 clone blocks across 7 Python version tables |
| Structural AST Search | 35 guard-and-raise statements in 8 files in 119.90 ms |
| Semantic Search | 10 hits in 8 ms across 3,133 declarations |
| Program Slicing | 99.8% reduction on can_be_split (135K lines down to 240 lines) |
| Predicate Inversion | 6 return statements inverted, not cancelled, non-call reference refusal proven |
| Parameter Object | (src, dst) bundled into @dataclass CodePair across 12 call sites in 4 files |
| RAM Pre-Validation | Injected attribute error caught in 0.68s, 8 baseline warnings subtracted |
| Remote Execution | 0.2s command execution in remote container |
Cite this article
Alexander Panasenko (2026-09-30). Black Under the Microscope: What 67 AST Tools Found Inside Python's Uncompromising Formatter. https://prod.codes/blog/black-under-the-microscope-67-ast-tools/