design note · · 5 min

From 1,623 µs to 129 µs: zero-copy AVX-512 wire streaming

A 5.7 MB binary wire block stream took 73 milliseconds to parse with standard MessagePack deserialization, stalling the ingest thread during market bursts. How an AVX-512 scanner and fixed-point integers dropped per-block latency from 1,623 µs to 129 µs and brought throughput to 978 MB/s.

On this page · 5 sections
  1. The Wire Format
  2. Why the Naive Parser Stalled
  3. The Zero-Copy SIMD Engine
  4. Hardware Benchmarks on Zen 4
  5. Production Impact

A market data engine connecting to a high-throughput financial L1 faces a stark choice. Public WebSockets either sample order books at 500 ms intervals or drop connections under volatility. Running a full follower node solves sampling, but demands dozens of gigabytes of RAM and saturates NVMe write endurance reconstructing state that trading bots discard within milliseconds.

The direct path is the exchange’s peer-to-peer gossip wire: raw TCP connections on the node communication ports, streaming consensus blocks as they are validated. You bypass public rate limits and receive unsampled executions tens of milliseconds ahead of public APIs.

There is one complication: the wire stream arrives as compressed, multi-megabyte binary MessagePack frames containing thousands of heterogeneous action payloads per block.

Our initial parser followed standard Rust practice: scan slices for action tags, deserialize them into serde_json::Value, and parse prices as floating-point numbers. On a 5.7 MB stream of 45 wire blocks, that parser took 73 milliseconds (1.62 ms per block), whereas the SIMD zero-copy engine finished in 5.8 milliseconds (129 µs per block). In volatile market bursts with ten blocks per second, the ingest thread saturated, backpressure mounted, and packets were dropped before they ever reached our matching engine.

Here is how replacing generic deserialization with an AVX-512 SIMD scanner, a zero-copy MessagePack cursor, and 10⁸ fixed-point integer arithmetic dropped per-block latency from 1,623 µs to 129 µs — a 12.5x speedup running at 978 MB/s with bit-for-bit data parity.


The Wire Format

Nodes communicate over TCP framing: a 4-byte big-endian payload length, a 1-byte compression flag (0x01 compressed with LZ4 or Zstandard, 0x00 raw), and the binary payload.

Once decompressed, the payload is an unindexed MessagePack stream containing consensus state, block headers, state hashes, and user actions. The actions appear as dynamically nested structures:

// Order action
["order", { "orders": [ [asset, is_buy, "146.10", "71.92", false, "Gtc", cloid] ], "grouping": "na" }]

// Cancel action
["cancel", [ [asset, oid], [asset, oid_2] ]]

// Batch modify action
["batchModify", { "modifies": [ [oid, [asset, is_buy, "146.20", "50.00", false, "Alo", cloid]] ] }]

Three characteristics shape the parsing problem:

  1. String-Encoded Financial Numbers: Prices and quantities are encoded as ASCII decimal strings (e.g. "146.10", "71.92").
  2. Heterogeneous Framing: Cancels appear as 2-element tuples [asset, oid], objects {"a": 0, "o": 12345}, or flat lists.
  3. No Top-Level Index: Action tags ("order", "cancel", "batchModify") are scattered within deeply nested maps and arrays.
Architecture diagram comparing legacy scalar search and heap deserialization against SIMD zero-copy parsing
Two paths from wire bytes to shared memory. The legacy path allocates hundreds of thousands of heap containers per second; the SIMD zero-copy path touches only contiguous memory.

Why the Naive Parser Stalled

The initial prototype parsed blocks with standard high-level abstractions:

// ❌ Scalar byte-by-byte substring scan
for window in data.windows(needle.len()) {
    if window == needle { /* found action */ }
}

// ❌ Deserializing untyped Value trees onto the heap
let val: serde_json::Value = rmp_serde::from_slice(&action_slice)?;

// ❌ Float parsing on the critical path
let px = ord["limit_px"].as_str().unwrap().parse::<f64>()?;
let sz = ord["sz"].as_str().unwrap().parse::<f64>()?;

Profiled on an AMD Ryzen 9 7950X3D, the bottleneck was immediate:

  1. Scalar Byte Scanning: windows(needle.len()) compiles to scalar comparison loops. Inspecting 5.7 MB of memory one byte at a time wastes the 512-bit vector registers sitting idle on Zen 4 cores.
  2. Heap Thrashing via serde_json::Value: Every parsed action allocated nested heap maps, string buffers, and variant vectors. A single 120 KB block caused thousands of small heap allocations. Memory bandwidth saturated and allocator locks dominated the profile.
  3. IEEE-754 Parsing Overhead: Calling str::parse::<f64>() performs mantissa accumulation, exponent scaling, and division according to IEEE-754 rounding rules. Financial systems immediately multiply this float back into a fixed-point integer, paying parsing overhead twice while risking precision drift.

The Zero-Copy SIMD Engine

We rebuilt the decoder under three strict rules:

  1. Zero heap allocations while traversing actions.
  2. AVX-512 hardware vectorization for needle searches.
  3. Branch-minimized fixed-point integer parsing for prices and quantities.

1. Vectorized Tag Discovery

Instead of a sliding byte window, we use memchr::memmem::Finder, which searches in 64-byte chunks using AVX-512 vector instructions (falling back to NEON on ARM):

use memchr::memmem::Finder;

let order_finder = Finder::new(b"order");
let cancel_finder = Finder::new(b"cancel");
let modify_finder = Finder::new(b"batchModify");

// Vectorized scan across the entire uncompressed block
let mut hits = Vec::with_capacity(512);
for pos in order_finder.find_iter(wire_bytes) {
    hits.push((pos, ActionTag::Order));
}

Once a tag candidate is identified, we probe backward by up to 16 bytes to locate the enclosing MessagePack array header (fixarray 0x90..=0x9f or array16 0xdc). This confirms the match is an authentic action envelope and not arbitrary string data inside a signature.

2. Zero-Allocation MessagePack Cursor

Rather than materializing intermediate structures, a custom cursor (FastRmpReader) advances directly through the slice:

pub struct FastRmpReader<'a> {
    data: &'a [u8],
    pos: usize,
}

impl<'a> FastRmpReader<'a> {
    #[inline(always)]
    pub fn read_u64(&mut self) -> Option<u64> {
        let b = *self.data.get(self.pos)?;
        self.pos += 1;
        match b {
            0x00..=0x7f => Some(b as u64),
            0xcc => {
                let v = *self.data.get(self.pos)?;
                self.pos += 1;
                Some(v as u64)
            }
            0xcd => {
                let slice = self.data.get(self.pos..self.pos + 2)?;
                self.pos += 2;
                Some(u16::from_be_bytes(slice.try_into().ok()?) as u64)
            }
            0xce => {
                let slice = self.data.get(self.pos..self.pos + 4)?;
                self.pos += 4;
                Some(u32::from_be_bytes(slice.try_into().ok()?) as u64)
            }
            0xcf => {
                let slice = self.data.get(self.pos..self.pos + 8)?;
                self.pos += 8;
                Some(u64::from_be_bytes(slice.try_into().ok()?))
            }
            _ => None,
        }
    }
}

Every reader method preserves positional rollback: if a field fails type verification, self.pos restores to its entry point. Speculative branches never leave the cursor in a corrupted state.

3. 10⁸ Fixed-Point Decimal Arithmetic

Exchange perpetuals standardize prices and lot sizes on a fixed factor of 10⁸ ($1.00000000 = 100,000,000$). Instead of converting ASCII strings to f64 and multiplying, we parse ASCII digits directly into a scaled u64:

#[inline(always)]
pub fn parse_fixed_point_8(bytes: &[u8]) -> Option<u64> {
    if bytes.is_empty() {
        return None;
    }
    let mut int_part: u64 = 0;
    let mut frac_part: u64 = 0;
    let mut in_frac = false;
    let mut frac_digits = 0usize;

    for &b in bytes {
        if b == b'.' {
            if in_frac { return None; }
            in_frac = true;
            continue;
        }
        if !b.is_ascii_digit() {
            return None;
        }
        let digit = (b - b'0') as u64;
        if !in_frac {
            int_part = int_part.checked_mul(10)?.checked_add(digit)?;
        } else if frac_digits < 8 {
            frac_part = frac_part * 10 + digit;
            frac_digits += 1;
        }
    }

    while frac_digits < 8 {
        frac_part *= 10;
        frac_digits += 1;
    }

    int_part.checked_mul(100_000_000)?.checked_add(frac_part)
}

"146.10" turns into 14,610,000,000 via three multiplications and an addition. Zero floating-point instructions are emitted, zero heap memory is touched, and rounding ambiguity is eliminated.


Hardware Benchmarks on Zen 4

We benchmarked the legacy parser against the zero-copy engine on production hardware:

  • CPU: AMD Ryzen 9 7950X3D (16 physical cores, 4.2 GHz base, 5.7 GHz boost, AVX-512)
  • Dataset: 45 real wire blocks (5.70 MB uncompressed MessagePack stream)
  • Measurement: 50 iterations (2,250 blocks, 766,900 actions decoded)
Benchmark comparison bar chart showing block latency and throughput improvements on Zen 4
Benchmark results on AMD Ryzen 9 7950X3D. Latency dropped from 1,623 µs to 129 µs; throughput reached 978 MB/s.
================================================================================
                    P2P WIRE PARSER HARDWARE BENCHMARK
================================================================================
Dataset:                 5.70 MB uncompressed wire blocks (45 blocks)
Total Processed:         2,250 blocks, 766,900 actions decoded
Parity Verification:     100.00% MATCH across all 766,900 actions
--------------------------------------------------------------------------------
Metric                     Legacy Decoder       SIMD Zero-Copy       Improvement
--------------------------------------------------------------------------------
Total Time (50 runs)        3,653.61 ms            291.53 ms           12.5x faster
Latency per Block           1,623.83 µs            129.57 µs           12.5x faster
Throughput                     78.05 MB/s          978.11 MB/s          12.5x throughput
Heap Allocation          Thousands per block     Zero (Zero-Copy)       Eliminated
Price/Size Parsing       f64 string-to-float     Fixed-Point 10^8 Int   Direct math
Search Algorithm         data.windows().pos()    memchr AVX-512         Vectorized
================================================================================

Across 766,900 decoded actions (orders, cancels, batch modifies), the zero-copy engine produced results bit-for-bit identical to the reference parser.


Production Impact

At 129 µs per block, wire parsing stops being a bottleneck.

Parsed actions feed directly into our lock-free shared memory publisher (ringfire), committing order book deltas into /dev/shm in 27–54 nanoseconds. Trading processes on the same machine consume the stream with zero system call overhead.

In a continuous 41-minute verification run against the exchange’s public WebSocket across ten major perpetual instruments:

  • Trade Parity: 100.00% exact 1-to-1 match (15,990 / 15,990 trades, 0 phantom trades, 0 drops).
  • Top-of-Book (BBO) Parity: 99.76% agreement (differences limited to sub-millisecond network transit order).
  • Process Footprint: Fixed at 122 MB of RSS, compared to 40+ GB for the official follower node.

When your wire format is binary, generic deserializers are convenient for test fixtures, but fatal on the ingest path. Moving from serde_json::Value to hardware SIMD searching and fixed-point integers reclaimed 1.5 milliseconds per block — turning a lagging consumer into a sub-millisecond market data feed.

Cite this article
Citation
Alexander Panasenko (2026-09-30). From 1,623 µs to 129 µs: zero-copy AVX-512 wire streaming. https://prod.codes/blog/from-1623-us-to-129-us-zero-copy-wire-streaming/