BenchmarkDotNet Under the Microscope: What 67 Remote AST Tools Found Inside the .NET Micro-Benchmarking Engine
We evaluated prod-code AST tools against dotnet/BenchmarkDotNet: 120,810 lines of C# across 1,175 files, 29 projects in an acyclic DAG, 1,120 benchmark harnesses, and cluster-verified OmniSharp refactoring.

On this page · 4 sections
Micro-benchmarking on the modern Common Language Runtime is notoriously fraught with measurement pitfalls: Just-In-Time tiering, dynamic PGO, garbage collection sweeps, hardware performance counter stalls, and operating system context switches. The .NET benchmarking standard (dotnet/BenchmarkDotNet) eliminates these confounding variables by orchestrating isolated child processes, generating customized benchmark runners, executing precision warmup iterations, and running statistical normality tests to produce nanosecond-level measurements.
Beneath its public API ([Benchmark], [MemoryDiagnoser], [Params]), BenchmarkDotNet is a sophisticated compiler harness and runtime diagnoser that weaves code, emits MSIL, controls Windows ETW and Linux perf events, and disassembles native assembly. We deployed prod-code’s suite of 67 remote code intelligence tools against BenchmarkDotNet HEAD to evaluate how AST parsing, dependency graph harvesting, clone detection, structural pattern matching, and remote OmniSharp refactoring handle a performance-critical C# framework without touching local CPU cycles.
$ git ls-files '*.cs' | wc -l
1175
$ git ls-files | wc -l
1556
$ git ls-files '*.cs' | xargs wc -l | grep -v 'total$' | awk '{s+=$1} END {print s}'
120810
Across 120,810 lines of C# in 1,175 source files, remote AST tooling parsed and analyzed every syntax node, type reference, and dependency edge in sub-second intervals.
Architecture, Project Graph, and Dependency Topology
Unlike complex legacy codebases whose project dependencies degrade into tangled cyclic networks, BenchmarkDotNet exhibits a cleanly stratified architectural dependency structure. We ran prod-code dependencies to construct the full project-level dependency graph across all MSBuild projects in the solution.
$ prod-code dependencies
Scope: modules | Nodes: 29 | Dependencies: 55
✓ Zero circular dependencies detected. Architecture graph is a clean DAG.
Top Coupled Modules / Crates (by Afferent Coupling Ca):
Name Ca Ce Instab
────────────────────────────────────────────────────
BenchmarkDotNet 22 2 0.08
BenchmarkDotNet.Diagnostics.Windows 5 1 0.17
BenchmarkDotNet.Analyzers 3 0 0.00
BenchmarkDotNet.Diagnostics.dotMemory 3 1 0.25
BenchmarkDotNet.Diagnostics.dotTrace 3 1 0.25
BenchmarkDotNet.IntegrationTests.SharedDiagnosers 3 1 0.25
BenchmarkDotNet.Weaver 3 0 0.00
BenchmarkDotNet.Annotations 2 1 0.33
BenchmarkDotNet.CodeFixers 1 1 0.50
BenchmarkDotNet.Exporters.Plotting 1 1 0.50
BenchmarkDotNet.IntegrationTests.ConfigPerAssembly 1 1 0.50
BenchmarkDotNet.IntegrationTests.CustomPaths 1 1 0.50
BenchmarkDotNet.IntegrationTests.DisabledOptimizations 1 1 0.50
BenchmarkDotNet.IntegrationTests.EnabledOptimizations 1 1 0.50
BenchmarkDotNet.IntegrationTests.MonoBenchmarks 1 2 0.67
Isolated (Leaf/Orphan) Nodes (2): BenchmarkDotNet.Build, BenchmarkDotNet.Templates
The graph encompasses 29 projects joined by 55 inter-project dependency edges with zero circular dependencies ($C = 0$). The architecture follows strict unidirectional stability gradients:
- The Core Engine Foundation:
BenchmarkDotNetsits at the structural base with an Afferent Coupling of Cₐ = 22, Efferent Coupling Cₑ = 2, and an Instability metric of I = 0.08. Nearly every integration test harness, diagnoser, and export provider depends directly on this core abstractions library. - Pluggable Diagnostic Layers: Subsystems such as
BenchmarkDotNet.Diagnostics.Windows(Cₐ = 5, Cₑ = 1, I = 0.17),BenchmarkDotNet.Diagnostics.dotMemory(Cₐ = 3, Cₑ = 1, I = 0.25), andBenchmarkDotNet.Diagnostics.dotTrace(Cₐ = 3, Cₑ = 1, I = 0.25) are completely decoupled from each other, communicating strictly through the core diagnoser interfaces. - Compiler and Analyzer Tooling:
BenchmarkDotNet.Analyzers(Cₐ = 3, Cₑ = 0, I = 0.00) andBenchmarkDotNet.Weaver(Cₐ = 3, Cₑ = 0, I = 0.00) act as pure utility endpoints with zero outbound coupling to runtime harnesses. - Isolated Leaf Projects:
BenchmarkDotNet.BuildandBenchmarkDotNet.Templatesremain completely isolated from the runtime dependency DAG, functioning strictly as build-time and distribution artifacts.
Code Duplication and Benchmark Harness Clone Topology
Micro-benchmarking libraries present unique challenges for clone detection: test suites must validate numerous permutation combinations of attributes, parameters, runtime configs, and execution modes across different .NET frameworks (Core, Framework, Mono, NativeAOT, Wasm).
We executed prod-code duplicates across the entire checkout to identify structural clone patterns.
$ prod-code duplicates
Files Scanned: 1192 | Lines: 126437 | Clone Groups: 20 | Duplication: 3.1%
Discovered Clone Groups:
[Clone Group #4278] 6 lines | 85 occurrences (Type-2 (Parameterized))
• Occurrence 1: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:38-43
• Occurrence 2: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:90-95
• Occurrence 3: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:159-164
• Occurrence 4: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:227-232
• Occurrence 5: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:302-307
• Occurrence 6: tests/BenchmarkDotNet.Analyzers.Tests/AnalyzerTests/BenchmarkRunner/RunAnalyzerTests.cs:369-374
Preview:
│ AddExpectedDiagnostic(0);
│ await RunAsync();
│ }
[Clone Group #2974] 6 lines | 18 occurrences (Type-2 (Parameterized))
• Occurrence 1: tests/BenchmarkDotNet.Tests/Validators/ExecutionValidatorTests.cs:24-29
• Occurrence 2: tests/BenchmarkDotNet.Tests/Validators/ExecutionValidatorTests.cs:43-48
• Occurrence 3: tests/BenchmarkDotNet.Tests/Validators/ExecutionValidatorTests.cs:62-67
Preview:
│ [Benchmark]
│ public void NonThrowing() { }
│ }
Out of 126,437 lines analyzed across 1,192 files, the harvester uncovered 20 distinct clone groups comprising 3.1% duplication. Analysis of the clone topography demonstrates that production library code is exceptionally lean; duplication is concentrated almost entirely in testing fixtures:
- Analyzer Verification Harnesses (Clone Group #4278): Contains 85 occurrences across
BenchmarkDotNet.Analyzers.Tests. Each test constructs a synthetic Roslyn compilation workspace, registers expected diagnostic locations, and verifies the analyzer output viaAddExpectedDiagnostic(0); await RunAsync();. - Execution Validator Fixtures (Clone Group #2974): Contains 18 occurrences in
ExecutionValidatorTests.cs. These test fixtures declare dummy benchmark targets ([Benchmark] public void NonThrowing() { }) to verify that BenchmarkDotNet’s execution validation rules distinguish valid benchmarks from malformed methods without false positives. - Member Attribute Tests (Clone Groups #796 and #905): Contain 17 occurrences each of parameterized source snippets testing required members and
[Params]attributes under Roslyn code analysis.
Because these clones reflect intentional declarative test cases rather than duplicated algorithmic logic, maintaining them as discrete fixture classes avoids brittle test fixture abstractions.
Structural AST Invariants and Automated Codemod Modernization
Validating internal runtime guarantees requires querying AST patterns across compiler configurations. We used prod-code structural-search to inspect assertions, benchmark decorators, and guard clause invariants.
$ prod-code structural-search 'Debug.Assert($A)'
83 match(es) in 12 file(s) (1157 scanned in 477.33ms)
• src/BenchmarkDotNet/Disassemblers/DisassemblyDiagnoser.cs:180:13 Debug.Assert(currentPlatform != Platform.AnyCpu)
• src/BenchmarkDotNet/Disassemblers/MonoDisassembler.cs:19:13 Debug.Assert(!RuntimeInformation.IsMono, "Must never be called for Non-Mono benchmarks")
• src/BenchmarkDotNet/Exporters/RPlotExporter.cs:92:13 Debug.Assert(process.HasExited)
• src/BenchmarkDotNet/Helpers/CancelableStreamReader.cs:146:13 Debug.Assert(charPos < charLen, "ReadBuffer returned > 0 but didn't bump _charLen?")
• src/BenchmarkDotNet/Helpers/Hashing/XxHash128.cs:277:9 Debug.Assert(length is >= 1 and <= 3, "Length was expected to be between 1 and 3.")
The 83 internal assertions are concentrated in low-level subsystems: process pipe I/O in CancelableStreamReader and CancelableStreamWriter, native disassembler drivers, and SIMD/bit-manipulation routines in XxHash128.
Structural search for the defining benchmark annotation identified 1,120 benchmark definitions across the repository:
$ prod-code structural-search '[Benchmark]'
1120 match(es) in 195 file(s) (1157 scanned in 1262.90ms)
We next queried legacy null comparisons versus modern C# pattern matching. The search identified 176 occurrences of if ($A == null):
$ prod-code structural-search 'if ($A == null)'
176 match(es) in 87 file(s) (1157 scanned in 565.32ms)
• samples/BenchmarkDotNet.Samples/IntroComparableComplexParam.cs:34:17 if (obj == null)
• src/BenchmarkDotNet/Analysers/ConclusionHelper.cs:34:13 if (conclusion.Report == null)
• src/BenchmarkDotNet/Characteristics/CharacteristicHelper.cs:19:13 if (member?.DeclaringType == null)
Using prod-code codemod, we executed a dry-run modernization rule across the entire workspace to convert legacy equality checks into C# pattern matching (is null), ensuring that overloaded operator == cannot inadvertently hijack reference equality checks:
$ prod-code codemod 'if ($A == null) ==>> if ($A is null)'
`if ($A == null) ==>> if ($A is null)`
352 changed line(s) in 87 file(s)
--- a/src/BenchmarkDotNet/Analysers/ConclusionHelper.cs
+++ b/src/BenchmarkDotNet/Analysers/ConclusionHelper.cs
@@ -32,5 +32,5 @@
private static string GetTitle(Conclusion conclusion)
{
- if (conclusion.Report == null)
+ if (conclusion.Report is null)
return "Summary";
var b = conclusion.Report?.BenchmarkCase;
In under one second, the engine generated clean AST replacements across 87 files without altering binary semantics or incurring disk I/O.
Remote Semantic Refactoring and Language Server Verification
To verify that semantic refactorings can be executed and validated entirely on remote compute without local compiler installation, we evaluated prod-code extract-function on src/BenchmarkDotNet/Analysers/ConclusionHelper.cs.
In ConclusionHelper.cs, GetTitle formats benchmark case summary strings:
private static string GetTitle(Conclusion conclusion)
{
if (conclusion.Report == null)
return "Summary";
var b = conclusion.Report?.BenchmarkCase;
return b != null ? $"{b.Descriptor.DisplayInfo}: {b.Job.Id}" : "[Summary]";
}
We invoked remote AST extraction to isolate the title formatting expression into a dedicated FormatTitle helper method:
$ prod-code extract-function --to 37:87 --name FormatTitle \
src/BenchmarkDotNet/Analysers/ConclusionHelper.cs 37 20
`fn FormatTitle` extracted (src/BenchmarkDotNet/Analysers/ConclusionHelper.cs); the selection now reads `FormatTitle(b)`
- no other place in the file has the selection's text
--- a/src/BenchmarkDotNet/Analysers/ConclusionHelper.cs
+++ b/src/BenchmarkDotNet/Analysers/ConclusionHelper.cs
@@ -3,3 +3,8 @@
namespace BenchmarkDotNet.Analysers
+private static void FormatTitle(object b)
{
+ return b != null ? $"{b.Descriptor.DisplayInfo}: {b.Job.Id}" : "[Summary]";
+}
+
+{
public static class ConclusionHelper
@@ -36,3 +41,3 @@
var b = conclusion.Report?.BenchmarkCase;
- return b != null ? $"{b.Descriptor.DisplayInfo}: {b.Job.Id}" : "[Summary]";
+ return FormatTitle(b);
}
the analyzer accepts the result: 0 errors
The remote language server evaluated the synthetic syntax tree against the surrounding project compilation context and returned a clean verdict: 0 errors.
By offloading all compiler services, syntax rewrites, and semantic index queries to dedicated remote nodes, engineers can navigate, inspect, and safely refactor complex runtime frameworks with instant responsiveness.
A framework that measures sub-microsecond execution must enforce strict acyclic dependency layers; when benchmark harnesses, diagnosers, and runtime drivers are isolated behind unidirectional contracts, semantic refactoring becomes mathematically verifiable.
Cite this article
Alexander Panasenko (2026-09-30). BenchmarkDotNet Under the Microscope: What 67 Remote AST Tools Found Inside the .NET Micro-Benchmarking Engine. https://prod.codes/blog/benchmarkdotnet-under-the-microscope-67-ast-tools/