From 1,623 µs to 129 µs: zero-copy AVX-512 wire streaming
A 5.7 MB binary wire block stream took 73 milliseconds to parse with standard MessagePack deserialization, stalling the ingest thread during market bursts. How an AVX-512 scanner and fixed-point integers dropped per-block latency from 1,623 µs to 129 µs and brought throughput to 978 MB/s.

On this page · 5 sections
A market data engine connecting to a high-throughput financial L1 faces a stark choice. Public WebSockets either sample order books at 500 ms intervals or drop connections under volatility. Running a full follower node solves sampling, but demands dozens of gigabytes of RAM and saturates NVMe write endurance reconstructing state that trading bots discard within milliseconds.
The direct path is the exchange’s peer-to-peer gossip wire: raw TCP connections on the node communication ports, streaming consensus blocks as they are validated. You bypass public rate limits and receive unsampled executions tens of milliseconds ahead of public APIs.
There is one complication: the wire stream arrives as compressed, multi-megabyte binary MessagePack frames containing thousands of heterogeneous action payloads per block.
Our initial parser followed standard Rust practice: scan slices for action tags, deserialize them into serde_json::Value, and parse prices as floating-point numbers. On a 5.7 MB stream of 45 wire blocks, that parser took 73 milliseconds (1.62 ms per block), whereas the SIMD zero-copy engine finished in 5.8 milliseconds (129 µs per block). In volatile market bursts with ten blocks per second, the ingest thread saturated, backpressure mounted, and packets were dropped before they ever reached our matching engine.
Here is how replacing generic deserialization with an AVX-512 SIMD scanner, a zero-copy MessagePack cursor, and 10⁸ fixed-point integer arithmetic dropped per-block latency from 1,623 µs to 129 µs — a 12.5x speedup running at 978 MB/s with bit-for-bit data parity.
The Wire Format
Nodes communicate over TCP framing: a 4-byte big-endian payload length, a 1-byte compression flag (0x01 compressed with LZ4 or Zstandard, 0x00 raw), and the binary payload.
Once decompressed, the payload is an unindexed MessagePack stream containing consensus state, block headers, state hashes, and user actions. The actions appear as dynamically nested structures:
// Order action
["order", { "orders": [ [asset, is_buy, "146.10", "71.92", false, "Gtc", cloid] ], "grouping": "na" }]
// Cancel action
["cancel", [ [asset, oid], [asset, oid_2] ]]
// Batch modify action
["batchModify", { "modifies": [ [oid, [asset, is_buy, "146.20", "50.00", false, "Alo", cloid]] ] }]
Three characteristics shape the parsing problem:
- String-Encoded Financial Numbers: Prices and quantities are encoded as ASCII decimal strings (e.g.
"146.10","71.92"). - Heterogeneous Framing: Cancels appear as 2-element tuples
[asset, oid], objects{"a": 0, "o": 12345}, or flat lists. - No Top-Level Index: Action tags (
"order","cancel","batchModify") are scattered within deeply nested maps and arrays.
Why the Naive Parser Stalled
The initial prototype parsed blocks with standard high-level abstractions:
// ❌ Scalar byte-by-byte substring scan
for window in data.windows(needle.len()) {
if window == needle { /* found action */ }
}
// ❌ Deserializing untyped Value trees onto the heap
let val: serde_json::Value = rmp_serde::from_slice(&action_slice)?;
// ❌ Float parsing on the critical path
let px = ord["limit_px"].as_str().unwrap().parse::<f64>()?;
let sz = ord["sz"].as_str().unwrap().parse::<f64>()?;
Profiled on an AMD Ryzen 9 7950X3D, the bottleneck was immediate:
- Scalar Byte Scanning:
windows(needle.len())compiles to scalar comparison loops. Inspecting 5.7 MB of memory one byte at a time wastes the 512-bit vector registers sitting idle on Zen 4 cores. - Heap Thrashing via
serde_json::Value: Every parsed action allocated nested heap maps, string buffers, and variant vectors. A single 120 KB block caused thousands of small heap allocations. Memory bandwidth saturated and allocator locks dominated the profile. - IEEE-754 Parsing Overhead: Calling
str::parse::<f64>()performs mantissa accumulation, exponent scaling, and division according to IEEE-754 rounding rules. Financial systems immediately multiply this float back into a fixed-point integer, paying parsing overhead twice while risking precision drift.
The Zero-Copy SIMD Engine
We rebuilt the decoder under three strict rules:
- Zero heap allocations while traversing actions.
- AVX-512 hardware vectorization for needle searches.
- Branch-minimized fixed-point integer parsing for prices and quantities.
1. Vectorized Tag Discovery
Instead of a sliding byte window, we use memchr::memmem::Finder, which searches in 64-byte chunks using AVX-512 vector instructions (falling back to NEON on ARM):
use memchr::memmem::Finder;
let order_finder = Finder::new(b"order");
let cancel_finder = Finder::new(b"cancel");
let modify_finder = Finder::new(b"batchModify");
// Vectorized scan across the entire uncompressed block
let mut hits = Vec::with_capacity(512);
for pos in order_finder.find_iter(wire_bytes) {
hits.push((pos, ActionTag::Order));
}
Once a tag candidate is identified, we probe backward by up to 16 bytes to locate the enclosing MessagePack array header (fixarray 0x90..=0x9f or array16 0xdc). This confirms the match is an authentic action envelope and not arbitrary string data inside a signature.
2. Zero-Allocation MessagePack Cursor
Rather than materializing intermediate structures, a custom cursor (FastRmpReader) advances directly through the slice:
pub struct FastRmpReader<'a> {
data: &'a [u8],
pos: usize,
}
impl<'a> FastRmpReader<'a> {
#[inline(always)]
pub fn read_u64(&mut self) -> Option<u64> {
let b = *self.data.get(self.pos)?;
self.pos += 1;
match b {
0x00..=0x7f => Some(b as u64),
0xcc => {
let v = *self.data.get(self.pos)?;
self.pos += 1;
Some(v as u64)
}
0xcd => {
let slice = self.data.get(self.pos..self.pos + 2)?;
self.pos += 2;
Some(u16::from_be_bytes(slice.try_into().ok()?) as u64)
}
0xce => {
let slice = self.data.get(self.pos..self.pos + 4)?;
self.pos += 4;
Some(u32::from_be_bytes(slice.try_into().ok()?) as u64)
}
0xcf => {
let slice = self.data.get(self.pos..self.pos + 8)?;
self.pos += 8;
Some(u64::from_be_bytes(slice.try_into().ok()?))
}
_ => None,
}
}
}
Every reader method preserves positional rollback: if a field fails type verification, self.pos restores to its entry point. Speculative branches never leave the cursor in a corrupted state.
3. 10⁸ Fixed-Point Decimal Arithmetic
Exchange perpetuals standardize prices and lot sizes on a fixed factor of 10⁸ ($1.00000000 = 100,000,000$). Instead of converting ASCII strings to f64 and multiplying, we parse ASCII digits directly into a scaled u64:
#[inline(always)]
pub fn parse_fixed_point_8(bytes: &[u8]) -> Option<u64> {
if bytes.is_empty() {
return None;
}
let mut int_part: u64 = 0;
let mut frac_part: u64 = 0;
let mut in_frac = false;
let mut frac_digits = 0usize;
for &b in bytes {
if b == b'.' {
if in_frac { return None; }
in_frac = true;
continue;
}
if !b.is_ascii_digit() {
return None;
}
let digit = (b - b'0') as u64;
if !in_frac {
int_part = int_part.checked_mul(10)?.checked_add(digit)?;
} else if frac_digits < 8 {
frac_part = frac_part * 10 + digit;
frac_digits += 1;
}
}
while frac_digits < 8 {
frac_part *= 10;
frac_digits += 1;
}
int_part.checked_mul(100_000_000)?.checked_add(frac_part)
}
"146.10" turns into 14,610,000,000 via three multiplications and an addition. Zero floating-point instructions are emitted, zero heap memory is touched, and rounding ambiguity is eliminated.
Hardware Benchmarks on Zen 4
We benchmarked the legacy parser against the zero-copy engine on production hardware:
- CPU: AMD Ryzen 9 7950X3D (16 physical cores, 4.2 GHz base, 5.7 GHz boost, AVX-512)
- Dataset: 45 real wire blocks (5.70 MB uncompressed MessagePack stream)
- Measurement: 50 iterations (2,250 blocks, 766,900 actions decoded)
================================================================================
P2P WIRE PARSER HARDWARE BENCHMARK
================================================================================
Dataset: 5.70 MB uncompressed wire blocks (45 blocks)
Total Processed: 2,250 blocks, 766,900 actions decoded
Parity Verification: 100.00% MATCH across all 766,900 actions
--------------------------------------------------------------------------------
Metric Legacy Decoder SIMD Zero-Copy Improvement
--------------------------------------------------------------------------------
Total Time (50 runs) 3,653.61 ms 291.53 ms 12.5x faster
Latency per Block 1,623.83 µs 129.57 µs 12.5x faster
Throughput 78.05 MB/s 978.11 MB/s 12.5x throughput
Heap Allocation Thousands per block Zero (Zero-Copy) Eliminated
Price/Size Parsing f64 string-to-float Fixed-Point 10^8 Int Direct math
Search Algorithm data.windows().pos() memchr AVX-512 Vectorized
================================================================================
Across 766,900 decoded actions (orders, cancels, batch modifies), the zero-copy engine produced results bit-for-bit identical to the reference parser.
Production Impact
At 129 µs per block, wire parsing stops being a bottleneck.
Parsed actions feed directly into our lock-free shared memory publisher (ringfire), committing order book deltas into /dev/shm in 27–54 nanoseconds. Trading processes on the same machine consume the stream with zero system call overhead.
In a continuous 41-minute verification run against the exchange’s public WebSocket across ten major perpetual instruments:
- Trade Parity: 100.00% exact 1-to-1 match (15,990 / 15,990 trades, 0 phantom trades, 0 drops).
- Top-of-Book (BBO) Parity: 99.76% agreement (differences limited to sub-millisecond network transit order).
- Process Footprint: Fixed at 122 MB of RSS, compared to 40+ GB for the official follower node.
When your wire format is binary, generic deserializers are convenient for test fixtures, but fatal on the ingest path. Moving from serde_json::Value to hardware SIMD searching and fixed-point integers reclaimed 1.5 milliseconds per block — turning a lagging consumer into a sub-millisecond market data feed.
Cite this article
Alexander Panasenko (2026-09-30). From 1,623 µs to 129 µs: zero-copy AVX-512 wire streaming. https://prod.codes/blog/from-1623-us-to-129-us-zero-copy-wire-streaming/