Every cell is a FIX round trip measured on the wire between two network namespaces. Every percentile and max is the worst reported value across all sessions.
These numbers are not publishable, and they are not
comparable with scan. The certified harness (bench/load/bench.sh) cannot run on this
box at all: every cell it launches is sudo -n ip netns exec … against root-owned artifacts under
/usr/local/libexec/libhft-bench, and it refuses to place work on non-isolated cores. What is tabulated
here came from the same binary (bench_roundtrip) with the same scenario flags and the same
recipe, driven directly over loopback. Four things make it indicative:
/sys/devices/system/cpu/isolated is empty on this host,
so the scheduler moved every server and client thread freely, across SMT siblings, for the whole run. The
machine file that makes placement explicit on scan has no counterpart here because there is
nothing to place onto.scan’s numbers cross a
real wire between two namespaces; a loopback round trip and a wire round trip are different measurements that
happen to share a unit.-march=native on this
box resolves to znver4 under gcc 15.2; the reference bench box builds
-march=broadwell, and the production fleet is on gcc 11.5. Newer codegen on
different silicon is a deliberate property of this host — it is the only lab box with
AVX‑512 VNNI — but it means a number here and a number there differ by toolchain
and ISA as well as by machine.Applies to Target throughput and the historical diagnostic tables. Every measured cell shows p50, p99, p99.9 and max for the selected latency. Service is actual client start to completed reply; response starts at the scheduled slot and includes scheduler delay. Missing values show —. All four acceptor models stay visible; highlighting compares the selected metric among visible, qualified results.
Temporary values — do not compare flavours on them. Cells marked smoke are behaviour probes, not steady-state latency measurements. The default smoke recipe is 200 warmup / 200 measured / one run, with one session for single and four for mono, pool and shard; the table headings describe the full 32-session recipe.
Use full-recipe measurements for latency comparisons: 20 000 warmup before each run, 20 000 measured, three runs, with the stability and parity gates satisfied.
not measured Highest sustained completed request/reply rate across the whole cell, established by an open-load rate sweep. Report the passing/failing rate bracket and response latency at the passing rate. This measures the configured client/server system, not an isolated engine limit.
No validated maximum-throughput sweep is available yet. The historical closed-loop test allows only one outstanding request per session and is not maximum throughput. Existing target-rate observations are not maximum-throughput results either. Candidate-run instructions and remaining validation work.
No target-throughput (open-loop) cell is measured on this host yet. The closed-loop matrix -- every measured cell so far -- is the expanded block below; the open sweep fills the table above as its cells complete.
These are separate experiments, retained for investigation rather than the two main throughput views. Their rates are not a backend-capacity measurement.
| Flavour | Transport | 1 · Single1 acceptor · 1 session | 2 · Mono1 acceptor thread · 32 sessions | 3 · PoolThread pool · 32 sessions | 4 · ShardSharded acceptors · 32 sessions |
|---|
| Flavour | Transport | 1 · Single1 acceptor · 1 session | 2 · Mono1 acceptor thread · 32 sessions | 3 · PoolThread pool · 32 sessions | 4 · ShardSharded acceptors · 32 sessions |
|---|
BENCH-CONTRACT §1 mode paced: each session is paced at the cell aggregate divided by its session count; the rate beneath each model's metrics is achieved/target and is flagged when the cell could not offer the load it was asked for. Above capacity a paced row is a saturation result, not a latency row.
| Flavour | Transport | 1 · Single1 acceptor · 1 session | 2 · Mono1 acceptor thread · 32 sessions | 3 · PoolThread pool · 32 sessions | 4 · ShardSharded acceptors · 32 sessions |
|---|
Use Latency to switch between Service (default) and Response. Each cell shows p50 / p99 / p99.9 / max in µs; Rate is achieved / target requests/s per session. Missing historical response metrics show —, never service values or zero. Rates and qualifications are unchanged by the selection.
Response p99.9 and max were recovered for all 59 retained measured OPEN cells from their original validated client logs. Existing values and verdicts are unchanged; this is publication repair, not a new benchmark. Recovery evidence.
BENCH-CONTRACT §1 mode open: fixed send schedule, independent of replies. How to run this bench. Response runs from the intended request-start slot to completed reply, including scheduler delay. Service runs from the client's actual request start to completed reply: the whole client/network/server exchange, not server processing alone. For each request, response = scheduler delay + service. The target is shared: 50 000/s for one session, or floor(50 000/32) = 1 562/s per session (49 984/s effective aggregate) at 32 sessions. Rate annotations are per session, using the slowest session's achieved rate. Blank response metrics mean unrecorded, never zero. Cells are replaced as the current sweep completes; see all current/retained v3 sitting tables documented in the handbook for progress and provenance. Cells listed in any of those tables have v3 evidence, subject to their reported status; invalidated observations are withdrawn, not usable measurements. C++/.NET open values that included warmup in the extraction sink were replaced by corrected reruns. Values absent from every table are historical. Retained baselines and newer sweep cells can use separate artifact revisions: combined coverage is not one identical-binary sitting. Closed and paced baselines are retained under Historical diagnostics, not relabelled as either throughput experiment.
Publication reconciled: retained sitting evidence restores
59 full-delivery cells: 10 ok and 49 partial / unstable.
Available partial timings are shown with an explicit qualifier and never receive best-result
highlighting. One genuine Java/JNI / VMA / pool failure remains; all 12 .NET Pure
cells are withdrawn for missing client CPU-pin application. See the
source-by-source reconciliation.
32-session pacing limitation: all three multi-session layouts currently align each process's eight session deadlines, producing batches roughly every 640.204 µs rather than uniformly spaced aggregate sends. Catch-up can add further bursts, and the four processes do not share a common schedule origin. The target and achieved rates are averages, not spacing checks. These v3 values describe that bursty workload; they must not be presented as measurements of a uniformly staggered stream. Offline replay of all 15 kernel flavour/layout combinations shows local median gaps of roughly 5–7 µs and p90 gaps of 585–601 µs, versus 80.026 µs for an evenly spaced client process. See the replay scope and limitations.
.NET integrity correction and pinning investigation: full request ownership and reply-stage guards now prevent duplicate or malformed replies from reusing timestamps. All 165 .NET tests pass. Existing latency rows predate this change; no old values are relabelled as corrected measurements. An installed eight-cell kernel smoke delivered fully but exposed missing CPU-pin application in the managed open/multi client. Archived-source proof confirms that all twelve managed OPEN latency/rate values require withdrawal; any retained numbers for those cells are invalid pending fresh pinned runs, not evidence of delivery failures. Their original full-delivery evidence is retained. Smoke timings are not imported; the pacing policy is unchanged. See verification details.
Fresh installed .NET verification is tracked after each kernel smoke cell; these short-run timings do not replace full-size target-throughput values.
Remaining open-load failure: Java/JNI/VMA pool completed 16/32 sessions in R10. Separate logging and authorized statistics diagnostics reproduced the failure; their timings are excluded. The statistics run completed 24/32 sessions and narrowed its failed client's deficit to transmit/progress before server consumption, without proving the cause. See the handbook for the retained evidence and limitations.
Recovery comparison completed: separate ten-second-grace and matched two-second-control JNI/VMA pool sittings both delivered 32/32 sessions, with stability qualified and no watchdog timeouts. Identical installed binaries were used for the pair. There is no evidence that longer patience fixed the intermittent stall; the canonical watchdog stays at two seconds. These diagnostics do not replace matrix values.
.NET Native shard correction: the R7 kernel/Onload shard numbers were withdrawn after logs and code showed only shard 0 was started. All 32 connections were served by that one acceptor; delivery counts alone did not validate the requested four-shard topology. All three .NET Native shard transports have now been rerun after repair. In R10 VMA, all four listeners started, but admission was 32/0/0/0: that result does not demonstrate four-way traffic distribution or sharding speedup.
Pinning-audit limitation: the earlier continuous watcher discarded an audit-script error, so empty watch logs do not prove clean runtime placement. Launcher/VMA affinity records remain available. Corrected sampling starts at 02:58:01 UTC during R7 Java/JNI/Onload/pool; earlier coverage is inconclusive. Low-duty activity on the core shared by the pool accept thread and worker 0 has been flagged. R10's watcher recorded 106 PASS, 24 VIOLATION and 13 INCONCLUSIVE samples; it does not establish uniformly clean placement. See the handbook for details.
| Acceptor model | p50 | p99 | p99.9 | max | sessions |
|---|---|---|---|---|---|
| thread-per-session (32 thr, 1 core) | 2 020 903 | 3 879 632 | 3 971 263 | 5 287 856 | 32/32 |
| 1 multiplexer | 219.40 | 1706.36 | 2179.62 | 4752.58 | 32/32 |
| 4-thread pool (least-loaded) | 100.97 | 1534.98 | 7064.26 | 10 050.52 | 32/32 |
A cell is the round-trip time of a FIX NewOrderSingle → ExecutionReport, measured on the wire between two network namespaces on one host. The clock starts at the client's scheduled send, not its actual send, so a client that falls behind is charged for the lateness instead of hiding it.
All of it is in one versioned file, load-bench-standards.conf, read
by the wrapper, by every harness and by the checker. The common recipe is
deliberately not per-scenario: the columns are comparable only if the work
per message is identical in each.
| common | value |
|---|---|
| warmup | 20 000, discarded |
| iterations | 20 000 measured |
| offered throughput | 50 000/s aggregate · fixed-rate open |
| jvm heap | 192 MB |
| scenario | nos_er |
| per scenario | sessions | acceptors | threads | runs |
|---|---|---|---|---|
| 1 single | 1 | 1 | 1 | 3 |
| 2 mono | 32 | 1 | 1 | 3 |
| 3 pool | 32 | 1 | 4 | 3 |
| 4 shard | 32 | 4 | 1 each | 3 |
The publication harness passes the complete recipe explicitly;
bare client or launcher invocations are diagnostics, not substitutes for the harness.
Every result carries bench_params (the argument vector),
bench_profile (what the server configured) and
bench_standards (the defaults in force, with a checksum of the file
that supplied them). check-bench-params.sh refuses a table whose
cells disagree, or whose profile contradicts the declared expectation.
Pinning policy is machine-independent — cores isolated, server and client sets disjoint, server from the front of the isolated set and clients from the back. The core numbers live in a per-host file, because a standard carrying one box's core list cannot be run on another, and a matrix that cannot be reproduced elsewhere is a single data point.
poll(2) (or zf_muxer on TCPDirect), drained from a
single poll snapshot before re-polling.SO_REUSEPORT, each on its own core, share-nothing. The
kernel decides where a connection lands — there is no assignment policy.
The standard uses four acceptors; M=1 is a separate diagnostic control.3 and 4 are not the same thing. The pool and shard columns are separate across all flavours; unsupported native-stack combinations remain explicitly marked no pool / no shard.
≥ is a saturated clamp, not a reading —
latency_u32() tops out at 232−1 ns.Both sides busy-poll, so a core carrying a server thread and a client starves both. Placement is therefore stated per scenario, not derived at run time.
No machine file for beelink2; the harnesses fall back to deriving cores from /sys/devices/system/cpu/isolated.
Server cores come from the front of the isolated set and client cores from the back, so widening the server — a pool, or more shards — eats into the middle and the two sets stay disjoint by construction rather than by arithmetic that has to be got right each time.
Unpinned is not neutral. isolcpus removes the
isolated cores from every process's default affinity mask, so a thread that is
never pinned does not land on a random isolated core — it lands on the
housekeeping cores beside every interrupt, while the cores reserved for
measurement sit idle. The JNI and .NET pool workers shipped that way for an hour;
it would have read as "the pool is slow".
Not every shared core is contention. java-pure pins its accept loop
and its multiplexer to the same core, and the accept loop blocks in
accept() — it burns nothing.
pinning-audit.sh --live samples each thread's CPU twice and reports a
collision only when more than one of them actually ran.
| flavour | 1 single | 2 mono | 3 pool | 4 shard |
|---|---|---|---|---|
| cpp | yes | yes | yes* | yes |
| java-pure | yes | yes | yes* | yes* |
| java-jni | yes | yes | yes* | yes |
| dotnet-native | yes | yes | yes* | yes |
| dotnet-pure | yes | yes*† | yes* | yes |
| rust-pure | yes | yes | yes | yes |
*newly capable, not yet measured. †dotnet-pure's mono-thread cells served a thread per session — it had no multiplexed mode until now — so they are not comparable with the rest of that column. Thread-per-session with cores to spare pays no multiplexing cost, so it reads as fast rather than as different: its kernel cell sat at 26.14µs against a single-session 25.61 for exactly that reason. java-pure's mono cells are multiplexed (p99 207, not the 81 002 of its thread-per-session default); it simply had to be asked, and now the standards ask.
Three flavours now have a true pool. C++ gained
FixAcceptorThreadPool — one listener, M workers, least-loaded
at accept, verified serving 8/8 sessions at 2 per worker — and
dotnet-pure gained the equivalent in its bench server.
java-jni reuses the very same serve pass the single-acceptor path runs, with accept disabled, so there is one serve implementation rather than two that drift; verified serving 8/8 sessions at p50 25.79µs.
rust-pure is the FIX engine written in Rust: its own client and server, no hftnet.
rust-native, the Rust client driving the C++ engine through the hftnet C ABI, was parked on
2026-09-11: its server was cpp's, so its multi-session cells re-measured the C++ server.
All six now pool on kernel TCP, Onload and VMA interposition. dotnet-native uses the
native FixAcceptorThreadPool, with per-worker connection ids, status,
errors and event rings projected through the existing managed session. Stack
listener pools (TCPDirect and SocketXtreme) remain unavailable.
TCPDirect has no reuse-port listener API. SocketXtreme accepts
independent SO_REUSEPORT binds, but four process pollers disconnect
every client during Logon because they cannot safely own the shared completion
stream. Both native-stack shard groups are marked no shard, not measured
as one acceptor under a shard label.
Latency comparisons only mean something if the cells do the same work. They do not, and it is invisible in the numbers. All four end up persisting nothing, by four different mechanisms at four different costs:
| flavour | acceptor | store | audit |
|---|---|---|---|
| cpp | FixAcceptor | absent (no template parameter) | absent |
| cpp (--acceptor-impl runner) | FixSessionRunnerAcceptor | null, compile-time | present, off |
| java-jni | FixSessionRunnerAcceptor | selected at runtime from config | present, off |
| dotnet-native | FixSessionRunnerAcceptor | branches on a null pointer per call | present, off |
| dotnet-pure | C# FixConnection | separate implementation | — |
All flavours now declare this at runtime. Each server
prints a bench_profile line stating the acceptor shape, store, audit,
schedule, TLS and thread count it actually configured, so the table above is
verified rather than read off the source, and
check-bench-params.sh asserts it against the declared expectation.
It earns its keep. Wiring java-pure's multiplexer, its thread count
was defaulted from the pool scenario's, quietly turning the mono-thread
cells into a 4-thread pool. Nothing in the latency showed it; the declaration read
threads=4 under a heading that says mono-thread, and that is how it
was caught.