design note · · 8 min

ringfire: the ring itself, on the other machine

A market-data ring in /dev/shm now has identical copies on other hosts, same sequence numbers, same order: 30 µs across a LAN by multicast, 51.7 ms p99 across the Pacific.

On this page · 4 sections
  1. What a mirror is
  2. Records by UDP, repairs by TCP
  3. What the stress test found
  4. What it costs

A market-data process on one machine parses exchange frames into a ring in /dev/shm. Strategies, testers and recorders on that machine read the ring at about 100 ns per hand-off and never see the network. Then the fleet grows: a 128-core box for testers, a cloud region next to an exchange, a second site. Every one of them wants the same stream, in the same order, as close to local as physics allows, and none of them may slow the source down.

We ruled out message brokers on day one: an extra process, an extra serialization, tens of microseconds per hop before any network, and a queue whose ordering and retention semantics are not the ring’s. What we wanted was simpler: the ring itself, on the other machine. That is what ringfire 0.5.0 ships as ringfire serve on the source host and ringfire mirror everywhere else.

What a mirror is

A ring is a fixed array of slots, each carrying a sequence number and a payload; for variable-length records the payload is a descriptor into a byte arena. The writer publishes sequence n into slot n mod capacity, and every reader keeps its own cursor. That structure is what gets copied.

The server maps the ring read-only and reads records exactly like a consumer does, slot protocol and all, but as raw bytes: it does not know the element type, so one serve/mirror pair works for fixed-size and blob rings alike. The mirror creates a ring from the geometry the source sends (capacity, slot size, schema signature, arena) and writes each record under its original sequence number with the same slot protocol writers use. Readers on the mirror host attach with RingConsumer or BlobConsumer, the same code as on the source host, and never touch the network. One flag differs: a mirror ring is sparse, because a mirror that joins mid-stream starts at the source’s current sequence. Consumers on a sparse ring skip to the next record present instead of waiting for one that will never come.

One record's path: producer pushes into the source ring, serve reads raw slot bytes, DATA frames cross the network, the mirror writes the same sequence number into an identical ring, readers attach as if local; 0.1 µs on the same ring, 3.8 µs to a mirror on the same host, 30 µs across a LAN, 51.7 ms from Tokyo to Los Angeles
One record's path. The mirror host's ring has the same geometry and the same sequence numbers; its readers cannot tell.

Records by UDP, repairs by TCP

Every mirror opens one TCP connection to the source: HELLO with the sequence it wants, GEOMETRY back, then either a stream of DATA frames on that connection or a MULTICAST frame saying where the datagrams will come from. On a LAN that is a multicast group and one sendto reaches every mirror behind the switch. Between sites it is 0.0.0.0, meaning unicast from one source socket; the mirror then sends PUNCH datagrams until data arrives, which opens its NAT from the inside. A frame is a 16-byte header plus raw slot bytes, up to 26 records in a 1500-byte datagram.

Order is the rule everything else follows. The mirror writes at exactly the next sequence, never elsewhere. A datagram that arrives early is held back. A hole is asked for with NAK on the TCP connection and answered from the source ring, which is the retransmission buffer: nothing is copied anywhere for later. A mirror may lag the source by up to capacity records before a range is overwritten (262,144 slots is 262 ms at 1 M msg/s, 26 s at 10 k msg/s); past that it receives GAP, and its readers see the same lapped count a slow reader on the source host would. Reordering and duplication cannot happen. What can happen, only under overload, is a visible gap.

Sequence diagram: HELLO, GEOMETRY and MULTICAST over TCP, PUNCH and DATA over UDP, a lost datagram, later data held back, NAK over TCP answered from the source ring, records written in order, GAP when the ring no longer holds the range, HEARTBEAT every millisecond while idle
The wire protocol. Live records travel by UDP; the handshake and every repair travel on the TCP connection each mirror already has.

Clouds do not route multicast: a VPC has none, and Transit Gateway multicast domains are a separate product. So a site runs one mirror over the WAN and serves it again locally. A mirror is an ordinary ring, and serve does not care who wrote it. One copy crosses the ocean however many readers the site has, and every hop keeps the source’s sequence numbers, so a leaf resumes, NAKs and gaps exactly like a direct mirror.

Source host with producer and ring sends one multicast datagram to a switch, which delivers it to mirror hosts 1, 2, 3 and 16, each running mirror, ring and readers
On a LAN: one datagram per frame whatever the number of mirrors. Sixteen mirrors: multicast p99 72 µs round trip, all complete; a TCP stream per mirror, p99 235 µs and two mirrors behind.
Tokyo source serves UDP unicast with every datagram sent twice across the internet to a hub mirror in AWS; the hub's ring is served again by unicast to three instances, each with mirror, ring and readers
Between sites: UDP unicast with NAT punching and duplicates over the WAN, then a hub mirror serves the instances of the site.

What the stress test found

An open-loop stress (publish at a fixed rate for a fixed time, never wait, match echoes by sequence) found three real bugs before any number was worth reporting.

The multicast heartbeat announced the ring’s write sequence. A record published between the sender’s last look at the ring and that load was announced before its datagram went out, and every mirror asked for it. Heartbeats now carry the sender’s last sent sequence.

Mirror::step drained every queued datagram before returning, so a single-threaded caller that interleaves its own work never got control back at 500 k msg/s. It now handles at most 32 datagrams per call.

A sender that keeps pace with the producer sends one record per datagram. Above about 20,000 msg/s that is a system call and a packet per record, and this kernel path sustains about 40,000 datagrams per second: latency went from 50 µs to over a millisecond. Frames are now paced. They leave at most once per 50 µs unless full, and a lone record arriving later than that after the previous frame goes out at once, so quiet and bursty streams pay nothing. Two other pacing rules were tried first. Waiting “whenever the previous frame was recent” delayed steady 20 k/s streams; waiting “whenever a backlog was seen” delayed the tails of bursts. Both were measured out.

Round trip p50 by publish rate over multicast: one record per datagram 51, 830 and 1200 µs at 20k, 50k and 100k messages per second; paced frames 88, 213 and 221 µs
Why frames are paced. Above the kernel's datagram rate the unpaced path queues up; paced frames stay near 200 µs round trip at 100,000 msg/s.

Two measurement artifacts are worth passing on. A mirror that connects with latest 130 ms after the producer started looks like a constant 2.5 % loss; it is history, not loss. And a mirror process given one core for its two busy-polling threads reports millisecond latencies quantized by the scheduler. Two cores per mirror.

What it costs

Two Ryzen 9 7950X hosts on a 1 GbE LAN, Linux 6.8, kernel network stack, 64-byte records, everything busy-polling. The master stamps each record on push; a consumer on the master and one on each of eight mirrors stamp the read, with mirror clocks translated into the master’s by a PTP-style offset from the minimum-round-trip probe.

Push-to-read p50 in microseconds: 0.1 on the same ring, 3.8 for a mirror on the same host by multicast, 9 by TCP, 10.3 by UDP unicast, 30 on another LAN host by multicast, 32 by TCP, 43 by UDP unicast, 52.9 for a leaf behind a hub
Push on the source to read by a consumer, p50 at 1,000 msg/s. The network hop is the cost; the mirror itself adds 3.8 µs.

Sustained, the multicast path delivered 100 % with zero reorders from 1,000 to 1,000,000 msg/s over 5 s runs; near 1 M/s the 1500-byte MTU, 26 records per datagram, is the limit. Across the Pacific, source in Tokyo and mirror in Los Angeles behind a home NAT at 1,000 msg/s, the shape of the tail is what differs.

Tokyo to Los Angeles one-way latency in milliseconds by percentile: TCP 50.4, 50.5, 99.6, 127.2, 132.2; UDP unicast 51.7, 51.7, 51.7, 56.2, 61.2; UDP unicast with every datagram twice 51.5, 51.5, 51.6, 53.6, 57.6
Tokyo to Los Angeles. TCP spends a full round trip recovering one record in a hundred; the UDP path's p99 sits 50 µs above its median, and sending twice trims the last of the tail.

Through a hub on the LAN a leaf saw 52.9 µs against 10.3 µs direct: the hub costs its two network hops and nothing measurable of its own. The recipes are short:

# source host, LAN: one datagram for all mirrors
ringfire serve /dev/shm/ticks --bind 0.0.0.0:7400 \
    --multicast 239.255.0.1:7401 --iface 10.0.0.5 --spin
ringfire mirror 10.0.0.5:7400 /dev/shm/ticks --iface 10.0.0.7 --spin

# source host, mirrors anywhere: unicast from port 7403, sent twice
ringfire serve /dev/shm/ticks --bind 0.0.0.0:7400 \
    --udp 7403 --dup 2 --spin
ringfire mirror source.example:7400 /dev/shm/ticks --unicast --spin

# a site hub: mirror the source, serve the mirror ring to the site
ringfire mirror source.example:7400 /dev/shm/ticks --unicast --spin &
ringfire serve /dev/shm/ticks --bind 0.0.0.0:7400 --udp 7403 --spin

The rule we would give anyone copying a ring across machines: do not put a queue between the ring and its copy. The ring is already the queue. It is the retention model, the retransmission buffer and the ordering guarantee, and every transport we tried was better for having nothing to add.

Cite this article
Citation
Alexander Panasenko (2026-09-26). ringfire: the ring itself, on the other machine. https://prod.codes/blog/ringfire-the-ring-itself-on-the-other-machine/