# FIX Benchmark Scenario Contract

This document is the canonical contract for the Java and C++ FIX benchmark
runners under:

- `bench/java/fixbench`
- `bench/cpp/fixbench`

The goal is to keep every engine on the same wire workflow. A benchmark row is
only comparable when both sides follow the same scenario, timestamp boundaries,
market-data depth, warmup exclusion, logging policy, and field-extraction work.

Before launching a cross-machine campaign, run the local preflight in
[`RUNBOOK.md`](RUNBOOK.md). It
checks the five libhft flavours against this scenario contract, verifies that
every registered external candidate is classified as covered, unsupported or
blocked for each canonical workflow, and documents how foreign engines with
non-Maven build systems are built from an external workspace instead of
vendoring source trees or build outputs into this repo.

## Global Rules

- All traffic is synthetic and local/cross-machine test traffic only. No broker,
  UAT, or production FIX endpoint is used.
- Warmup samples are discarded. Raw samples, CDFs, and summary percentiles use
  measured samples only.
- Per-message FIX logging is disabled for headline rows. Optional audit traces
  may log a small sample of generated messages before measurement starts.
- Validation, stores, dictionaries, and debug callbacks are disabled unless the
  row is explicitly labelled as a correctness or diagnostic row.
- `|` in examples below represents FIX SOH (`0x01`).
- Standard market-data scenarios use `md.levels=10`: request `264=10`, snapshot
  `35=W`, `268=20` bid/ask entries.
- Tick-to-trade uses top-of-book only: `ttt.md.levels=1`, `268=2` bid/ask
  entries (`35=W` snapshot ticks for `tick_to_trade`/`risk_maker`, `35=X`
  incremental ticks for `stream_ttt`).
- Every outbound response or reaction is causally derived from the parsed
  inbound message that triggered it. A benchmark implementation must not use a
  fixed prebuilt reply/order except for immutable header/template fragments.
- Payload content varies deterministically across measured messages. The
  variation is part of the benchmark contract: it prevents one-message constant
  replay, unrealistic branch prediction, dead-code elimination, and overfitting
  to a single tag/value layout.
- Session layer stays realistic: `108=HeartBtInt` is 30 seconds for every row,
  logon uses `141=ResetSeqNumFlag=Y`, and admin messages (heartbeats, test
  requests) remain enabled and are processed normally during measurement.
  Samples that happen to interleave with admin traffic are kept, for every
  engine alike, and report-grade rows record the count of admin messages
  processed inside the measured window (per direction) so a tail spike can be
  attributed to a heartbeat rather than to the engine path. Any `35=2`
  ResendRequest observed inside the measured window invalidates the row.
- TCP framing: `TCP_NODELAY` on both ends. Headline rows write one FIX message
  per send call; application-level batching of scenario messages is only
  allowed in explicitly burst-labelled rows. Receivers must tolerate arbitrary
  TCP segmentation of inbound messages; a partial-read path being exercised
  does not invalidate a sample.

## Load Model

The harness paces sends on a fixed-interval schedule derived from the row's
throughput target. Two properties of that loop are part of the contract and
must be disclosed with every report-grade row:

- `t0` is stamped at the actual task start, not the intended send time. When an
  iteration overruns its slot, the schedule is re-anchored to now. Recorded
  latencies are therefore service-time samples: queueing delay caused by the
  engine falling behind the offered rate is not folded into the latency
  distribution (coordinated omission is not corrected).
- To make that visible instead of silent, report-grade rows must record the
  scheduler-delay probe (`taskStart - intendedSendTime`) and report achieved
  throughput next to the target. A row whose achieved throughput is below 98%
  of target, or whose scheduler-delay p99 exceeds the send interval, is
  saturated: it must be reported as a saturation/throughput result, never as a
  headline latency row.

Rates are chosen per run profile (see Run Profiles). Comparisons are only valid
between rows with the same target rate and the same wait strategy.

## Active Scenario Names

| Scenario | Flow | Primary sample owner | Purpose |
|---|---|---|---|
| `nos_er` | `35=D -> 35=8` | Client | Order submit and execution-report round trip. |
| `md` | `35=V -> 35=W` | Client | Market-data request and snapshot round trip. |
| `md_nos_er` | `35=V -> 35=W -> 35=D -> 35=8` | Client | Reactive order flow with parsed MD feeding the order. This replaces the old mixed/business rows. |
| `tick_to_trade` | `35=W -> 35=D` | Server/feed | Completed tick send to reaction-order arrival. This replaces the old `marketmaker` name. |
| `risk_feed` | `35=V -> 35=W` | Client | Feed migration profile; same wire contract as `md`, consumed via the risk-app feed topology/profile. |
| `risk_order` | `35=D -> 35=8` | Client | Order migration profile; same wire contract as `nos_er`, consumed via the risk-app order topology/profile. |
| `risk_maker` | `35=W -> 35=D` | Server/feed | Maker migration profile; same wire contract as `tick_to_trade`, consumed via the risk-app maker topology/profile. |

### The canonical `35=8` — pinned, because "same field set" was not true

Every arm answering `35=D` MUST emit this ExecutionReport. It is pinned by VALUE and by
FIELD SET, not by "a valid FIX 4.4 ER", because two arms were shipping different messages
while both passing a field-set check:

| Tag | Field | Value | Note |
|---|---|---|---|
| 37 | OrderID | arm-defined | |
| 17 | ExecID | arm-defined | |
| **150** | **ExecType** | **`0` (New)** | an ACK, the first response a venue sends |
| **39** | **OrdStatus** | **`0` (New)** | must agree with the quantities below |
| 11 | ClOrdID | echoed from the `35=D` | |
| 55 | Symbol | echoed | |
| 54 | Side | echoed | |
| 38 | OrderQty | echoed | |
| 14 | CumQty | `0` | |
| 151 | LeavesQty | `= OrderQty` | the whole order rests |
| 44 | Price | echoed | |
| 6 | AvgPx | `0` | |
| 60 | TransactTime | arm-defined | |

**ExecType is `0`, never `2` or `1`.** FIX 4.4 REMOVED `ExecType` FILL(`2`) and
PARTIAL_FILL(`1`); a 4.4 fill is `150=F` (Trade) with `39=2`. An ack is the correct reply
here in any case: the fill is a later event, and order-entry latency is measured to the
first response.

**LastQty (32) and LastPx (31) are deliberately ABSENT.** They are fill-only fields —
"Conditional (fill)" in FIX 4.4 — and an ack has traded nothing, so carrying them at
`0` is not a richer message, it is a wrong one. libhft was emitting both and now does not.

This reverses an earlier decision recorded here, and the reason it reversed is worth
keeping. The plan had been to add those two fields to the NexusFix arm rather than remove
them from libhft's, on the principle that trimming your own arm to do less work is
indistinguishable from cheating. That principle stands — but it does not apply here,
because the fields do not belong on this message in the first place. The evidence settled
it: NexusFix's FIX 4.4 encoder emits LastQty only `if (last_qty_.raw > 0)`, and LastPx only
alongside it, so adding them at `0` would have changed nothing on the wire while looking
like it had. A no-op that reads as a fix is worse than either option.

Net effect on libhft's message: minus `LastQty`/`LastPx`, plus `TransactTime` — one field
fewer, for FIX correctness rather than for the number. Stated plainly because it does move
libhft's encoding cost down slightly, and that must not be discovered rather than declared.

An earlier note also claimed libhft emitted five fields NexusFix lacked. That was read off
NexusFix's extraction *sink* rather than its message builder and was wrong; the builder
already carried `ClOrdID`, `OrderQty` and `Price`.


## Current Engine Scenario Coverage

This section keeps the active coverage table for the whole benchmark universe:
internal controls, commercial/runtime comparison engines, open-source
full-session candidates, and parser-only references. The open-source discovery
lane is also exported by
`/mnt/c/Temp/libhft/open-source-fix-engines/results/scenario_coverage.csv`;
the combined repo-side inventory is
[`COVERAGE.csv`](COVERAGE.csv).

Status shorthand:

- `S5 pass`: stable cross-machine Stage 5 evidence.
- `S5 local pass`: stable two-process local TCP evidence.
- `S5 local diag` / `S5 diag`: runnable, but not report-grade yet.
- `S5 failed`: a cross-machine Stage 5 attempt was made but failed; this is
  visible evidence of the current blocker, not coverage.
- `S4 diag`: local workflow evidence only.
- `S3 smoke`: same-process smoke only.
- `preflight pass`: local source/build/contract proof that the row is wired for
  the canonical workflow, but not yet report-grade latency evidence.
- `alias:*`: behavior-preserving legacy wire name, usually `marketmaker` for
  `tick_to_trade`.
- `approx:*`: legacy approximation, not parity-comparable for the canonical
  flow.

`canon` describes the scenario shape. It does not, by itself, mean the row is
headline parity-grade. `S5 diag canon` is the deliberate label for rows whose
wire scenario is canonical but whose current evidence is missing some local
pre-campaign proof such as an observable extraction sink, `contract_status=pass`,
or the revised tick-to-trade measurement boundary.

### Combined Coverage At A Glance

| Engine | Lane | Language | Build | Covered | Weakest | `nos_er` | 10-level `md` | `md_nos_er` | `tick_to_trade` |
|---|---|---:|---|---:|---|---|---|---|---|
| `tcp-ping-pong-cpp` | network baseline | C++ | build-ok | n/a | baseline RTT only | n/a | n/a | n/a | n/a |
| `tcp-ping-pong-java` | network baseline | Java | build-ok | n/a | baseline RTT only | n/a | n/a | n/a | n/a |
| `tcp-ping-pong-net` | network baseline | C#/.NET | build-ok | n/a | baseline RTT only | n/a | n/a | n/a | n/a |
| `tcp-ping-pong-rust` | network baseline | Rust | build-ok | n/a | baseline RTT only | n/a | n/a | n/a | n/a |
| `libhft-cpp` | core control | C++ | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-java-jni` | core control | Java/JNI+C++ | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-java-pure` | core control | Java | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-dotnet-managed` | core control | C#/.NET Pure | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-dotnet-native` | core control | C#/.NET Native/hftnet | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-rust-pure` | core control | Rust | build-ok | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |
| `crossfix` | commercial/runtime | Java/native | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `philadelphia` | commercial/runtime | Java | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `philadelphia-fast` | commercial/runtime | Java | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `quickfixj` | commercial/runtime baseline | Java | build-ok in FixBench | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `artio` | open-source discovery | Java | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `dfx` | open-source discovery | Rust | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `falcon` | open-source discovery | Java | build-ok | 1/4 | S3 smoke | S3 smoke canon | unsupported | unsupported | unsupported |
| `fix8` | open-source discovery | C++ | build-ok via bench3-compatible bundle | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `fixantenna_net` | open-source discovery | C# | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `fixer_rs` | open-source discovery | Rust | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `libtrading` | open-source discovery | C | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `llfix` | open-source discovery | C++ | adapter-ready; pinned source and retained evidence required | 0/4 | none | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing |
| `nanofix` | open-source discovery | Rust | build-ok with pinned Aeron overlay | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |
| `nexusfix` | open-source discovery | C++ | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `openfix` | open-source discovery | C++ | tooling-blocked | 0/4 | none | blocked:tooling-blocked | blocked:tooling-blocked | blocked:tooling-blocked | blocked:tooling-blocked |
| `quickfix_cpp` | open-source discovery | C++ | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `quickfix_go` | open-source discovery | Go | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `quickfix_n` | open-source discovery | C# | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `robaho_cpp_fix_engine` | open-source discovery | C++ | compile-failed | 0/4 | none | blocked:compile-failed | blocked:compile-failed | blocked:compile-failed | blocked:compile-failed |
| `truefix` | open-source discovery | Rust | build-ok with V/W dictionary overlay | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |
| `ferrumfix` | parser/codec reference | Rust | compile-failed upstream / stage2-shim-ok | n/a | parser-only | n/a | n/a | n/a | n/a |
| `fixpp` | parser/codec reference | C++ | build-ok | n/a | parser-only | n/a | n/a | n/a | n/a |
| `hffix` | parser/codec reference | C++ | build-ok | n/a | parser-only | n/a | n/a | n/a | n/a |
| `robaho_cpp_fix_codec` | parser/codec reference | C++ | build-ok | n/a | codec-only | n/a | n/a | n/a | n/a |

Fix8 caveat: the original local Ubuntu-built shared libraries were not usable
on the SITEA GLIBC 2.34 hosts. The current Stage5 row uses a bench3-built
GLIBC-compatible Fix8/Poco bundle staged through `--fix8-compatible-root`; that
bundle now produces report-grade bench1/bench2 rows for all four canonical
workflows.

### Core FixBench Lane

These rows are the internal control engines and should be present in every
headline comparison. They are maintained in the FixBench result bundles rather
than the open-source discovery `scenario_coverage.csv`.

The six active libhft flavours are deliberately separate rows. C++, both Java
providers and both .NET providers have report-grade S5 rows. Rust Pure has
canonical four-flow localhost Stage 4 diagnostics; physical integration is
pending. Fresh campaigns must carry the same p50/p99/p99.9, GC/runtime and
pinning metadata before promotion.

| Engine | Language | Build | Covered | Weakest | `nos_er` | 10-level `md` | `md_nos_er` | `tick_to_trade` |
|---|---:|---|---:|---|---|---|---|---|
| `libhft-cpp` | C++ | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-java-jni` | Java/JNI+C++ | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-java-pure` | Java | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-dotnet-managed` | C#/.NET Pure | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-dotnet-native` | C#/.NET Native/hftnet | build-ok | 4/4 | S5 pass | S5 pass canon | S5 pass canon | S5 pass canon | S5 pass canon |
| `libhft-rust-pure` | Rust | build-ok | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |

### Open-Source Discovery Lane

| Engine | Language | Build | Covered | Weakest | `nos_er` | 10-level `md` | `md_nos_er` | `tick_to_trade` |
|---|---:|---|---:|---|---|---|---|---|
| `artio` | Java | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `dfx` | Rust | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `falcon` | Java | build-ok | 1/4 | S3 smoke | S3 smoke canon | unsupported | unsupported | unsupported |
| `fix8` | C++ | build-ok via bench3-compatible bundle | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `fixantenna_net` | C# | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `fixer_rs` | Rust | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `libtrading` | C | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `llfix` | C++ | adapter-ready; pinned source and retained evidence required | 0/4 | none | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing | blocked:retained-session-evidence-missing |
| `nanofix` | Rust | build-ok with pinned Aeron overlay | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |
| `nexusfix` | C++ | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `openfix` | C++ | tooling-blocked | 0/4 | none | blocked:tooling-blocked | blocked:tooling-blocked | blocked:tooling-blocked | blocked:tooling-blocked |
| `quickfix_cpp` | C++ | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag non-parity | S5 diag non-parity | S5 diag non-parity |
| `quickfix_go` | Go | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `quickfix_n` | C# | build-ok | 4/4 | S5 diag | S5 diag canon | S5 diag canon | S5 diag canon | S5 diag canon |
| `robaho_cpp_fix_engine` | C++ | compile-failed | 0/4 | none | blocked:compile-failed | blocked:compile-failed | blocked:compile-failed | blocked:compile-failed |
| `truefix` | Rust | build-ok with V/W dictionary overlay | 4/4 | S4 diag | S4 diag canon | S4 diag canon | S4 diag canon | S4 diag canon |

Next coverage work:

- Park fringe engine promotion unless a specific comparison need appears:
  `falcon` has no market-data support, and `openfix` /
  `robaho_cpp_fix_engine` remain tooling or portability blocked. `fixer_rs`
  now has bench1/bench2 Stage5 rows via a bench3-built GLIBC-compatible Rust
  adapter binary.
- Keep any future hardened reruns on the canonical scenario names. The current
  core C++/Java, C#/.NET, Java runtime comparison, and Artio diagnostic rows now
  have canonical `md_nos_er` and `tick_to_trade` evidence in the generated
  coverage table.
- **`nexusfix` Stage5 was hardened on 2026-08-13** and is the first
  open-source arm to be gated on its extraction sink. It now folds the canonical
  sink over the full required field set before `t1`, derives every payload from
  `bench/cpp/fixbench/src/wire_fixture.h`, echoes `262`/`55`/`264` into the
  `35=W` and `11`/`55`/`54`/`38`/`44` into the `35=8`, runs the server-owned
  `tick_to_trade` row through the shared `hft-bench` runner, prints
  `extraction_sink=` on both owner rows, and ships an in-process `--self-test`
  that the Stage-5 driver runs as a build gate. Its active source also pins the
  TTT timestamp boundary (`t0` before `35=W` send, `t1` at inbound `35=D`
  callback entry before extraction), emits no headline `35=8`, and makes the
  passive client prove the reactive order send completed. Its current
  checked-in coverage cells are deliberately `S5 diag canon`, not `S5 pass
  canon`, until fresh report-grade Stage-5 artifacts replace the older
  evidence; those fresh artifacts are only promoted when the owner rows carry
  the hardened runtime contract marker and a non-zero extraction sink.
- `quickfix_cpp` and `quickfix_n` Stage5 adapter **source now emits
  `extraction_sink=`** on the owner rows: QuickFIX C++ folds the received
  reaction `35=D`, `35=W`, and `35=8` fields through the shared C++ sink, and
  now carries the hardened runtime contract marker so fresh sink-gated artifacts
  can promote it. QuickFIX/N does the equivalent field fold in the active
  `bench/dotnet/fixbench/quickfix-n` project with the Stage-5 summary carrying
  the sink into CSV/report output, but still needs a runtime contract marker.
  Their checked-in cells are still diagnostic until fresh artifacts publish
  the missing evidence.
- QuickFIX C++'s active Stage5 source also now makes `tick_to_trade`
  server-owned with `t0` immediately before the completed `35=W` send, `t1` at
  inbound `35=D` callback entry before extraction, no `35=8` leg, and a passive
  client wait for the reactive order send to complete before shutdown. This is
  source/static-regression evidence; fresh Stage5 artifacts are still required.
- QuickFIX/N's active Stage5 source also now makes `tick_to_trade`
  server-owned with `t0` immediately before the completed `35=W` send, `t1` at
  inbound `35=D` callback entry before extraction, and no `35=8` leg. This is a
  source/static-regression fix; a fresh .NET 10 build and Stage5 run is still
  required before promoting its historical rows.
- `libtrading`, `quickfix_cpp`, and `quickfix_n` Stage5 adapters have partial
  canonical-wire hardening in the repo code: 10-level `md` / `md_nos_er`,
  TOB-only `tick_to_trade`, and no ExecutionReport leg in `tick_to_trade`.
  Their checked-in cells remain diagnostic or non-parity until fresh artifacts
  publish the missing observable sink / `contract_status=pass` evidence.
- Keep blocked engines visible rather than dropping them from the matrix.
  In particular, `llfix` is pinned by
  `tools/check-llfix-benchmark-readiness`; its canonical Stage-5 adapter and
  driver local gate are complete, but it remains `0/4` until retained pinned
  session, store/recovery and isolated-host evidence exists.
- Treat Artio and NexusFix's retained `S5 diag` cells as historical diagnostic
  coverage, not a comparison result. Their next-run source pins and named
  BLOCKED evidence requirements are checked by
  `tools/check-oss-comparator-benchmark-readiness`.

## Control Scenario Names

Control rows are not FIX engine scenarios, but they are part of the benchmark
contract and should be collected beside report-grade FIX rows.

| Scenario | Flow | Primary sample owner | Purpose |
|---|---|---|---|
| `baseline_rtt` | raw TCP echo | Client | Wire/kernel/NIC round-trip floor at matched transport and payload size. Control row for every report-grade table. |

## Planned Scenario Names

These flows are part of the benchmark contract but are not wired across the
active Java and C++ runners yet. They must not be reported as active headline
rows until both sides implement the same timestamp and extraction rules.

| Scenario | Flow | Primary sample owner | Purpose |
|---|---|---|---|
| `cxl_replace` | `35=G -> 35=8` | Client | Cancel/replace round trip against rolling order state. Hottest production order path for maker/risk flows. |
| `stream_ttt` | `35=V`, then `35=X ... -> 35=D` | Server | Streaming incremental-feed profile: subscribe once, react to marked paced/burst ticks. |

The `risk_*` profiles share their base scenario's wire contract and flow; they
differ in client consumption topology and session profile (owner-loop/batch
event consumption, feed/order ultra-low-latency profiles), which is an
engine-configuration difference, not a scenario difference. Message payload
for every scenario is selected by the pinned wire-shape fixture
(`fixbench.wire.shape`):

- `java` (default): the field sets defined in this document.
- `legacy`: risk-app payload. Adds to `35=D`: `1=Account`,
  `200=MaturityMonthYear`, `167=SecurityType`, `48=SecurityID`. Uses the
  legacy-style `35=8`: legacy `17=ExecID`/`37=OrderID` format, `39=0`, and the
  legacy `6=AvgPx`/`14=CumQty` constants.

Internal libhft control rows (`libhft-cpp`, `libhft-java-jni`,
`libhft-java-pure`, `libhft-dotnet-managed`, and `libhft-dotnet-native`) are
headline rows only when they use their ULL profiles: direct fixed-buffer C++,
pure managed Java/.NET, JNI-backed Java, hftnet/native, validation/store/FIX
logging off, pinned owner loops where applicable, compact event views where the
scenario can carry the required correlation marker, and no per-message object
materialisation. Compatibility paths such as libhft-java native-worker plus
Java poller are diagnostic rows and must be labelled separately.

The normative fixture constants (symbol, base prices, quantities, legacy field
values) live in the shared fixture
`bench/java/fixbench/src/main/java/org/latency/bench/FixBenchWire.java` and
its C++ counterpart; adapters must consume them from there, not duplicate
literals. The audit pack's captured sample messages are the record of what was
actually sent; a row whose captured payload differs across engines is not
comparable and must not be reported as one.

Note that `md_nos_er` spans two request/reply round trips plus derivation work.
Report it in its own table or clearly labelled; never alongside single-RTT rows
as if directly comparable.

Retired mixed names `nos_md_er`, `md_er`, `business`, and `nos_md_er_business`
are not accepted by active runners. `marketmaker` is retired in favour of
`tick_to_trade` but is deliberately kept as an **accepted alias** so existing
runner scripts keep working; new benchmark runs should use `tick_to_trade`.
Report loaders may map historical result directories, but new runs should use
the canonical names above.

## Timestamp Boundaries

### Client Request -> Reply Scenarios

For active client request/reply scenarios (`nos_er`, `md`, `md_nos_er`,
`risk_feed`, and `risk_order`):

- `t0` is captured on the client immediately before the engine-specific call
  that builds/encodes and writes the first outbound FIX message for the
  scenario.
- `t1` is captured on the client only after the final inbound message is fully
  read, parsed, and the required field extraction/checks have completed.

This means the measured latency includes client encoding, client socket write,
network, server read/parse, server response encoding/write, network, client
read/parse, and required application extraction/check work.

### Server Tick-To-Trade Scenarios

For `tick_to_trade`, `stream_ttt`, and `risk_maker`:

- The server/feed side first builds and finalises the tick message (`35=W`, or
  `35=X` for `stream_ttt`). Tick construction is outside the sample.
- `t0` is captured on the server/feed side immediately before the
  engine-specific call that writes the completed tick message.
- `t1` is captured on the same server/feed side at the first instruction in the
  benchmark application callback/handler after the engine has recognized and
  delivered the inbound `35=D`. It must precede benchmark-side field extraction,
  marker matching, business checks, or response generation. For a raw adapter,
  the equivalent boundary is immediately after framing and message-type
  recognition. Engine-internal framing and parsing are therefore included;
  benchmark application processing is not.
- After `t1`, the benchmark matches the `11=ClOrdID` tick marker and performs
  the required extraction/hash checks. A failed check discards the sample and
  fails the row; post-`t1` validation is never included in the latency value.

This measures the outbound feed send, feed delivery, client
receive/parse/decision/order build and send, the return network hop, and
server-side order framing and engine parsing. It deliberately excludes
server-side tick generation and benchmark-side order extraction/checking so
that it remains distinct from the fully processed request/reply scenarios.

### Baseline Round Trip

For `baseline_rtt` there is no FIX engine in the loop. The client stamps `t0`
into the payload immediately before the socket write; `t1` is captured on the
client after the echoed payload has been read back and the touch load (if any)
has completed. Both timestamps are on the same clock, so no cross-machine sync
is required.

## Required Field Extraction

Benchmarks must not satisfy a scenario by sending constant messages through a
socket and recording the reply. Every parse-oriented scenario performs a small,
deterministic amount of application work.

For every scenario, the triggered outbound message must use values extracted
from the inbound message:

- `nos_er`: server parses `35=D` and uses its `11`, `55`, `54`, `38`, `44`, and
  marker fields when building `35=8`.
- `cxl_replace`: client builds each `35=G` from its own rolling order state
  (`41=OrigClOrdID` is the previously accepted `11`); server parses `35=G` and
  uses its `11`, `41`, `55`, `54`, `38`, `44` when building the replace `35=8`.
- `md`: server parses `35=V` and uses its `262`, `55`, `264`, and requested
  entry types when building `35=W`.
- `md_nos_er`: client parses `35=W` and uses its marker, symbol, touch prices,
  and touch sizes when building `35=D`; server then parses that `35=D` when
  building `35=8`.
- `tick_to_trade` / `stream_ttt`: client parses the tick (`35=W` or `35=X`) and
  uses its marker, symbol, touch prices, and touch sizes when building `35=D`.

Static message templates are allowed only for the parts that are invariant in a
real engine, such as fixed tag order, FIX version, sender/target IDs, and
constant configuration fields. Mutable fields must be overwritten from the
current parsed inbound state before send.

### Extraction Sink — what it does and does not guarantee

⚠️ **The sink is a regression check, not an adversarial one.** It reliably
catches an engine that silently stops doing work between builds — the common
failure. It cannot distinguish a lazy consumer from a degenerate producer, and
it is forgeable. Two concrete attacks:

1. **Precomputation.** Payload variation is deterministic and the final sink is
   engine-independent for a given fixture/scenario/iteration count, so the
   expected value is knowable in advance. A consumer can hardcode the constant,
   skip parsing entirely and pass. *The determinism that makes it a cross-engine
   anchor is what makes it forgeable.*
2. **Producer-chosen values.** If the producer emits fields that collapse the
   fold — all-zero integers give `sink = sink*31 + 0`, which stays 0 — the sink
   stops discriminating, and a broken consumer becomes indistinguishable from a
   degenerate producer. The fold is also linear and public, so a passing value
   can be constructed without parsing.

Hardening, in order of value:

- **Per-run nonce** carried in the payload and folded. The sink can no longer be
  precomputed and changes every run, so a hardcoded constant fails immediately.
  Report the nonce with the result so runs stay reproducible on demand.
- **Per-message verification** rather than a single final fold (the `9001`/`9002`
  check tags below are the existing mechanism). A final fold cannot catch a
  consumer that parses one message and skips the remaining 4,999.
- **Vary payloads within a run**, so no single field value can dominate or
  collapse the fold.

This matters more as these benchmarks are used competitively: the incentive to
look fast is real, and a check that implies a guarantee it does not provide is
worse than a check known to be weak.

### Extraction Sink

Extraction must be observable, or the compiler/JIT is entitled to delete it.
Every parse/extract row — cross-machine scenario or microbenchmark — folds
every extracted field into a per-run checksum sink. This is universal, not
limited to rows labelled as business checks:

- The fold operation is pinned: `sink = sink * 31 + value` over unsigned
  64-bit values.
- Canonical value forms are pinned so no engine pays a representation
  conversion designed for another architecture: integer-valued fields
  (quantities, numeric ids) fold as int64; prices fold as scaled int64 at the
  fixture's pinned scale; single-char fields (`54`, `150`, `39`, `269`) fold
  as their ASCII byte; string fields (`55`, string ids) fold as the sum of
  their ASCII bytes.
- The final sink value is written to the run's result metadata.
- Because payload variation is deterministic, the final sink for a given
  fixture, scenario, and iteration count is engine-independent across
  **full-message-parsing engines** (QuickFIX/J, Philadelphia, CrossFIX, Artio),
  which can extract every listed field. Engines in that class reporting
  different sink values for the same row did not do the same work: the row is
  invalid, not comparable. This is the work-parity check.
- **Decoded-event engines** (e.g. libhft, which delivers a decoded event that
  surfaces a single `quantity`/`price` and per-entry ids rather than the
  distinct wire tags 38/151/14/6/44 and 262/268) fold the same canonical forms
  over the field set their model exposes. Their sink is engine-local by design
  and is **not** a cross-engine parity anchor; the anchor is the ExecutionReport
  sink of the full-message-parsing engines. A decoded-event engine that cannot
  visit a required field marks the row per the closing rule of this document
  rather than reporting it as parity-comparable.
- If the fixture includes benchmark check tags such as `9001`/`9002`, the
  implementation additionally verifies them per message and fails the row on
  mismatch.

## Deterministic Payload Variation

Payload variation should be deterministic and reproducible from a sequence
number, request id, or tick id. It should not require random allocation or wall
clock randomness inside the measured loop.

Minimum variation by scenario:

- `nos_er`:
  - `11=ClOrdID` changes every message.
  - `54=Side` alternates or follows a deterministic function of `ClOrdID`.
  - `38=OrderQty` and `44=Price` follow the pinned generator-side table walks
    (see Canonical Derivation Rules).
  - `35=8` echoes/derives those changing fields.
- `md`:
  - `262=MDReqID` changes every message.
  - `55=Symbol` rotates through a small configured symbol set when enabled.
  - bid/ask prices and sizes vary by level and request id while preserving
    `ask >= bid`.
  - `268` and level count match the configured depth.
- `md_nos_er`:
  - the `35=V`/`35=W` variation above applies.
  - the client derives `35=D` side, price, and quantity from the current parsed
    `35=W`, not from a constant order fixture.
  - the final `35=8` derives from that `35=D`.
- `tick_to_trade` / `stream_ttt`:
  - each tick has a fresh marker in `262=MDReqID`.
  - top-of-book bid/ask and sizes vary by tick.
  - the client derives `35=D` side, price, quantity, symbol, and `11=ClOrdID`
    from the current parsed tick.
- `cxl_replace`:
  - `11=ClOrdID` changes every message; `41=OrigClOrdID` chains to the previous
    accepted order.
  - `44=Price` and `38=OrderQty` follow the pinned generator-side table walks
    so each replace actually changes the order.

Varied numeric fields must vary in digit width as well as value: price and
quantity tables must include entries with different integer/fraction digit
counts, and id fields must grow through at least two widths during a measured
run. Constant-width variation still lets a parser specialize on one layout.

The variation should remain small enough that all engines exercise the same
business logic, but large enough that a benchmark cannot be optimized by
hardcoding one literal FIX payload or one branch path.

### MarketDataSnapshotFullRefresh (`35=W`) / IncrementalRefresh (`35=X`)

Every engine implementation must extract at least (for `35=X`, the same
logical fields from its `279`-tagged entries):

- `262=MDReqID`, used as the scenario marker or request/tick id.
- `55=Symbol`.
- `268=NoMDEntries`.
- For top-of-book:
  - bid entry: `269=0`, `270=MDEntryPx`, `271=MDEntrySize`
  - ask entry: `269=1`, `270=MDEntryPx`, `271=MDEntrySize`
- For 10-level MD rows, all bid/ask entries are scanned. Extraction can be
  direct ordered scanning, indexed lookup, or typed view access, but it must
  visit the same logical fields.

All extracted values fold into the extraction sink (see Extraction Sink
above). Business-check rows additionally verify per-message invariants
(`ask >= bid`, entry count matches `268`) and fail the row on violation.

### ExecutionReport (`35=8`)

Every engine implementation must extract at least:

- `11=ClOrdID`
- `37=OrderID`
- `17=ExecID`
- `55=Symbol`
- `54=Side`
- `150=ExecType`
- `39=OrdStatus`
- `38=OrderQty`
- `151=LeavesQty`
- `14=CumQty`
- `6=AvgPx`
- `44=Price` when present
- `41=OrigClOrdID` when present (`cxl_replace` rows)

All extracted values fold into the extraction sink before `t1` is recorded.

#### `150=ExecType` must not be `'2'` on a FIX 4.4 session

**FIX 4.4 REMOVED `ExecType='2'` (FILL) and `'1'` (PARTIAL_FILL).** A 4.4 fill is
`150=F` (Trade), with the fill state carried by `39=OrdStatus` (`'2'` Filled,
`'1'` Partially filled — those OrdStatus values are unchanged and remain valid).

This is not pedantry. A dictionary-validating engine session-rejects the message
(`35=3`, `371=150`, `373=5`) **before** it reaches `fromApp`, so the responder
believes it replied, the client counts no reports, and the arm reports a latency
built from whatever did arrive. The QuickFIX C++ arm failed exactly this way and
took eight rounds to characterise, because every layer reported success.

Every responder in this repo emitted an invalid value at some point:

| emitter | was | now |
|---|---|---|
| C++ `bench_roundtrip.cpp` (2 sites) | `'2'` | `'F'` |
| Java/JNI `hftjava_jni.cpp` (2 sites) | `'2'` | `'F'` |
| Java Pure `PureJavaSession` acceptor replies | `'2'` | `'F'` |
| .NET `hftnet_session.cpp` | `'2'` | `'F'` |
| Java bench `FixBenchWire.EXEC_TYPE_TRADE` | **`'2'`** | `'F'` |
| QuickFIX C++ stage4 + stage5 | `ExecType_FILL` | `ExecType_TRADE` |
| QuickFIX/N bench + stage3 + stage5 | `ExecType.FILL` | `ExecType.TRADE` |

The Java one is worth singling out: the constant was **named**
`EXEC_TYPE_TRADE` and **valued** `'2'`. Both Java responders read it, so the
entire Java side shipped an invalid ExecType from one character that the
identifier actively disguised. Grep for the tag, not for the name.

**Legitimately still `'2'`:** the FIX **4.2** dictionary validators
(`libhft/fix/generated/Dictionary42Validators.h`) and the Falcon stage-3 smoke harness, which
negotiates `8=FIX.4.2` where FILL is correct. Inbound parsers also still map
`'2'`→`Fill` — accepting a legacy value on receive is not the same as emitting
one.

### Canonical Derivation Rules

Wherever a scenario says a reactive order field is "deterministic" or
"derived" from the parsed tick/book state, the formulas are fixed. Leaving
them as prose is implementation room, and implementation room becomes
workload drift.

Side:

- `54=1` (buy) when the numeric marker id parsed from `262=MDReqID` is odd,
- `54=2` (sell) when it is even.

Price:

- buy: `44 = bestAsk` extracted from the tick,
- sell: `44 = bestBid` extracted from the tick.

Quantity:

- `38 = max(lotSize, min(maxQty, touchSize))`, where `touchSize` is the
  extracted size at the selected touch (ask size for buys, bid size for
  sells), and `lotSize`/`maxQty` are pinned by the fixture.

Generator-side varied fields (`nos_er`, `cxl_replace` requests, and tick
prices/sizes) follow pinned table walks:

- `44 = basePrice + (id % priceSteps) * tickSize`
- `38 = qtyTable[id % len(qtyTable)]`

with the tables chosen to satisfy the digit-width variation rule. The
constants (`basePrice`, `tickSize`, `priceSteps`, `qtyTable`, `lotSize`,
`maxQty`) are pinned by the fixture, identically for every engine.

The rule is deliberately trivial, but it must be computed from the parsed
marker at reaction time, not precomputed on the generator side, so every
engine pays the same parse-decide-derive path. Implementations must all use
this rule; a fixture may override it only if the override is applied to every
engine in the comparison.

## Scenario Details

### `nos_er`: NewOrderSingle -> ExecutionReport

Flow:

```text
Client t0
Client -> Server: 35=D NewOrderSingle
Server parses D
Server -> Client: 35=8 ExecutionReport
Client parses/extracts 8
Client t1
```

Client `35=D` must include at least:

- `11=ClOrdID`
- `55=Symbol`
- `54=Side`
- `38=OrderQty`
- `40=OrdType`
- `44=Price` for limit-style rows
- `60=TransactTime`

Server `35=8` must echo or derive enough fields for the client checks above.

Example:

```text
8=FIX.4.4|35=D|49=CLIENT|56=SERVER|34=2|52=20260707-08:00:00.000|11=1000001|55=EUR/USD|54=1|38=1000000|40=2|44=1.08501|60=20260707-08:00:00.000|10=...|
8=FIX.4.4|35=8|49=SERVER|56=CLIENT|34=2|52=20260707-08:00:00.001|37=O1000001|17=E1000001|11=1000001|55=EUR/USD|54=1|150=0|39=0|38=1000000|151=1000000|14=0|6=0|44=1.08501|10=...|
```

### `md`: MarketDataRequest -> MarketDataSnapshot

Flow:

```text
Client t0
Client -> Server: 35=V MarketDataRequest
Server parses V
Server -> Client: 35=W MarketDataSnapshotFullRefresh
Client parses/extracts W and runs MD check
Client t1
```

The default row uses `md.levels=10`, so `35=W` has `268=20` entries. Top-of-book
rows are only used for the tick-to-trade scenarios. Entry ordering is pinned:
levels run
best to worst, with the bid entry (`269=0`) immediately followed by the ask
entry (`269=1`) at each level, as in the example. Engines may exploit that
ordering, but the fixture must emit it identically for every engine.

Example:

```text
8=FIX.4.4|35=V|49=CLIENT|56=SERVER|34=2|52=20260707-08:00:00.000|262=2000001|263=0|264=10|267=2|269=0|269=1|146=1|55=EUR/USD|10=...|
8=FIX.4.4|35=W|49=SERVER|56=CLIENT|34=2|52=20260707-08:00:00.001|262=2000001|55=EUR/USD|268=20|269=0|270=1.08500|271=1000000|290=1|269=1|270=1.08502|271=1000000|290=1|...|10=...|
```

### `md_nos_er`: MarketDataRequest -> MarketDataSnapshot -> NewOrderSingle -> ExecutionReport

This is the canonical mixed/business scenario. It replaces the old
`nos_md_er`, `md_er`, `business`, and `nos_md_er_business` rows.

Flow:

```text
Client t0
Client -> Server: 35=V MarketDataRequest
Server parses V
Server -> Client: 35=W MarketDataSnapshotFullRefresh
Client parses/extracts W and runs MD check
Client builds 35=D from extracted W values
Client -> Server: 35=D NewOrderSingle
Server parses D
Server -> Client: 35=8 ExecutionReport
Client parses/extracts 8 and runs ER check
Client t1
```

The `35=D` in this scenario must be populated from the parsed `35=W`; it must
not be a prebuilt constant order. Minimum derivation:

- `11=ClOrdID` contains the `262=MDReqID` marker from the `35=W` so the response
  can be tied back to the request.
- `55=Symbol` is copied from `35=W`.
- Best bid/ask prices and sizes are extracted from `35=W`.
- `54=Side` is deterministic from the extracted tick/book state.
- `44=Price` is taken from the extracted touch:
  - buy order: use best ask
  - sell order: use best bid
- `38=OrderQty` is derived from the extracted touch size, capped by the row's
  configured order quantity.
- `60=TransactTime` is generated on order send.

Side, price, and quantity follow the Canonical Derivation Rules above: they
must depend on the parsed marker/book state and be identical across engine
implementations.

Example:

```text
8=FIX.4.4|35=V|...|262=3000001|264=10|267=2|269=0|269=1|146=1|55=EUR/USD|10=...|
8=FIX.4.4|35=W|...|262=3000001|55=EUR/USD|268=20|269=0|270=1.08500|271=1000000|290=1|269=1|270=1.08502|271=1000000|290=1|...|10=...|
8=FIX.4.4|35=D|...|11=3000001|55=EUR/USD|54=1|38=1000000|40=2|44=1.08502|60=20260707-08:00:00.002|10=...|
8=FIX.4.4|35=8|...|37=O3000001|17=E3000001|11=3000001|55=EUR/USD|54=1|150=0|39=0|38=1000000|151=1000000|14=0|6=0|44=1.08502|10=...|
```

### `tick_to_trade`: MarketDataSnapshot -> NewOrderSingle

This is the server-side tick-to-trade benchmark. It replaces the old
`marketmaker` name.

Flow:

```text
Server/feed builds top-of-book W
Server/feed t0   (captured immediately BEFORE sending the completed W)
Server/feed -> Client: 35=W MarketDataSnapshotFullRefresh
Client parses/extracts W
Client updates strategy/order state
Client builds 35=D from extracted W values
Client -> Server/feed: 35=D NewOrderSingle
Server/feed engine recognizes D and enters the application callback/handler
Server/feed t1   (first callback instruction, BEFORE benchmark field extraction)
Server/feed matches the tick marker and validates the extraction sink
```

The primary sample is `server_tick_to_trade = t1 - t0`. It includes the W send
call but excludes W construction and benchmark-side D extraction. Client-side
probes such as client parse time, order-round-trip, or tick-to-fill are useful
diagnostics but are not the headline tick-to-trade metric.

`35=W` is top-of-book only:

- `268=2`
- one bid entry
- one ask entry

The client must parse/extract the `35=W` and use it to populate the `35=D`.
Minimum derivation is the same as `md_nos_er`:

- `11=ClOrdID` carries the tick id from `262=MDReqID`.
- `55=Symbol` is copied from `35=W`.
- `44=Price` is best ask for buy or best bid for sell.
- `38=OrderQty` is derived from the selected touch size and configured cap.
- `54=Side` is deterministic from the parsed tick/book state.

For headline rows the server must not send a `35=8` in response to the `35=D`:
a reply on the measured socket is extra traffic that varies by implementation
choice. A tick-to-fill diagnostic row may enable the `35=8`, but then it must
be enabled for every engine in that comparison, and the row is labelled
`tick_to_fill`, not `tick_to_trade`.

Example:

```text
8=FIX.4.4|35=W|49=SERVER|56=CLIENT|34=2|52=20260707-08:00:00.000|262=4000001|55=EUR/USD|268=2|269=0|270=1.08500|271=1000000|290=1|269=1|270=1.08502|271=1000000|290=1|10=...|
8=FIX.4.4|35=D|49=CLIENT|56=SERVER|34=2|52=20260707-08:00:00.001|11=4000001|55=EUR/USD|54=1|38=1000000|40=2|44=1.08502|60=20260707-08:00:00.001|10=...|
```

### `baseline_rtt`: Raw TCP Round-Trip Floor

Reference implementations:

- C++: `bench/cpp/fixbench/src/ping_pong.cpp`
- Java: `bench/java/fixbench/src/main/java/org/latency/net/TcpPingPong.java`
- .NET: `bench/dotnet/fixbench/tcp-ping-pong/Program.cs`

No FIX engine is involved. The client writes a payload of `slots` u64 words,
each holding `t0`; the server echoes it; the client reads `t0` back out of the
payload and records `now - t0` on its own clock. All three runtimes emit the
same CSV columns and hft-bench-compatible raw binary distribution format. Two
variants:

- `--touch 0` (zero-touch): the server blind-echoes. Pure wire/kernel/NIC
  floor.
- `--touch 1` (random-access touch): both ends read a masked pseudo-random
  slot, forcing a real load of the received bytes the way a parser does. This
  is the floor to quote next to FIX rows.

Rules:

- Transport parity with the FIX row being decomposed: same plain-TCP or Onload
  transport, same `TCP_NODELAY`/`TCP_QUICKACK`, busy-spin receive, and pinning
  strategy.
- `slots` is chosen so the payload byte size approximates the FIX scenario's
  larger direction (e.g. the `35=W` size for MD rows, the `35=D`/`35=8` size
  for order rows).
- Every report-grade table includes the matched `baseline_rtt` rows. The
  decomposition `engine_cost = fix_rtt - baseline_rtt(touch 1)` is a
  first-order approximation, valid for means and medians only: the two runs
  are not sample-paired, and percentiles do not subtract. Never derive tail
  claims (p99 and beyond) by subtraction; quote the baseline's tail
  percentiles alongside the FIX row's tails instead.

### `cxl_replace`: OrderCancelReplaceRequest -> ExecutionReport

Setup (unmeasured, during warmup): the client establishes one resting order per
symbol with a `35=D -> 35=8` exchange, seeding the rolling order state.

Measured flow:

```text
Client t0
Client -> Server: 35=G OrderCancelReplaceRequest
Server parses G
Server -> Client: 35=8 ExecutionReport (150=5 replaced)
Client parses/extracts 8 (including 41) and updates rolling order state
Client t1
```

Client `35=G` must include at least:

- `11=ClOrdID` (fresh every message)
- `41=OrigClOrdID` (the previously accepted `11`)
- `37=OrderID` when the previous `35=8` carried one
- `55=Symbol`
- `54=Side`
- `38=OrderQty` (varied so the replace is real)
- `40=OrdType`
- `44=Price` (varied so the replace is real)
- `60=TransactTime` generated on send

Server `35=8` responds with `150=5`, `39=0`, echoing/deriving `11`, `41`, `37`,
and the changed price/quantity. The client must verify the `41` chain before
recording `t1`; a broken chain fails the row.

Example:

```text
8=FIX.4.4|35=G|49=CLIENT|56=SERVER|34=3|52=20260707-08:00:00.002|11=5000002|41=5000001|37=O5000001|55=EUR/USD|54=1|38=1500000|40=2|44=1.08503|60=20260707-08:00:00.002|10=...|
8=FIX.4.4|35=8|49=SERVER|56=CLIENT|34=3|52=20260707-08:00:00.003|37=O5000001|17=E5000002|11=5000002|41=5000001|55=EUR/USD|54=1|150=5|39=0|38=1500000|151=1500000|14=0|6=0|44=1.08503|10=...|
```

### `stream_ttt`: Subscribe Once, React to Streamed Incremental Ticks

This is the streaming-feed profile. `md` and `risk_feed` measure a
request/reply exchange per sample, which is a parse benchmark, not how a
production feed behaves: real consumers subscribe once and then process a
paced stream of incremental updates. `stream_ttt` covers that shape with the
same single-clock server-side measurement as `tick_to_trade`.

Flow:

```text
Client -> Server: one 35=V subscription (unmeasured)
Server -> Client: initial 35=W snapshot (unmeasured)
loop at stream.rate:
  Server/feed builds 35=X incremental update
  Server/feed t0 (marked ticks only)
  Server/feed -> Client: 35=X
  Client parses/extracts X, updates top-of-book state
  On marked ticks: client builds 35=D from current parsed book state
  Client -> Server/feed: 35=D
  Server/feed engine delivers D, t1
  Server/feed matches tick marker and validates D after t1
```

- `35=X` carries `262=MDReqID` as the tick marker and `268=2` entries
  (`279=0|1`, `269=0|1`, `55`, `270`, `271`), varying per tick per the payload
  variation rules.
- Every `stream.mark_every`-th tick (default 1) is marked; the client must
  react to marked ticks with a `35=D` derived per the `tick_to_trade` rules.
  Unmarked ticks must still be fully parsed and applied to book state.
- `stream.burst` (default 1) sends that many back-to-back updates in one
  socket write, with the marked tick last. Burst rows measure drain behavior:
  `t1 - t0` then includes the client working through the whole burst before
  reacting. Burst rows are labelled `stream_ttt_burst<N>` and are the only
  rows exempt from the one-message-per-send framing rule.
- The primary sample is `server_tick_to_trade = t1 - t0` for marked ticks.

If an engine cannot parse `35=X`, the row is marked unsupported per the
closing rule of this document; do not silently substitute `35=W`.

## Microbenchmarks

Component microbenchmarks should use the same fixture semantics as the
cross-machine scenarios:

- `parse_only` is parse plus light realistic extraction. It is not a pure SOH
  scan.
- `parse_extract` is parse plus full realistic extraction and hash/check.
- Market-data parsing should include `35=W`. Report-grade headline component
  rows use the same configured depth as the workflow `md` / `md_nos_er` rows:
  currently 10 levels, i.e. `268=20` bid/ask entries. `parse_only` reads only
  the top-of-book fields and verifies tag `9001`; `parse_extract` reads every
  `269`/`270`/`271` entry and verifies tag `9002`.
- Execution-report parsing should include `35=8`.
- Encoding tests should validate the generated buffer outside the timed hot
  path or with a timed checksum/hash if the row is explicitly an encode-check
  row.

Microbenchmarks follow the same Extraction Sink rule as the cross-machine
scenarios: every extracted field folds into the pinned per-run sink, the final
sink value is written to the result metadata, and fixture check tags are
verified when present. A microbenchmark whose sink diverges from the other
engines' on the same fixture is invalid.

Open-source parser discovery rows may use a smaller synthetic fixture while an
adapter is being brought up. Those rows must stay labelled as discovery evidence
and must not be ranked in the same report-grade component leaderboard as the
depth-10 headline rows until regenerated with the headline fixture contract.

## Run Profiles

Every report-grade row belongs to exactly one profile. Rows from different
profiles never share a table without labels.

### `headline`

Unsaturated latency at a fixed moderate rate.

- At least 5 measured runs per row; the first run after process start is
  discarded (JIT/warmup run).
- Java: at least 100k warmup iterations per scenario per fork before the
  measured phase. C++: at least 50k.
- At least 100k measured samples per run. A percentile `p` may only be quoted
  when the run holds at least `10 / (1 - p)` samples (p99.9 needs >= 10k,
  p99.99 needs >= 100k).
- Validation, stores, and dictionaries off, per the Global Rules.

### `load_sweep`

The same scenario at at least 3 target rates spanning the intended production
rate (suggested: 1k, 10k, 50k msg/s and upward until saturation). Each rate is
its own row. Report p50/p99/p99.9 versus rate plus the maximum sustained rate
(the highest target where achieved throughput stays >= 98% of target). This is
the only place a saturated row may appear, reported as throughput, not
latency.

### `soak`

At least 30 minutes sustained at a moderate rate. Report percentile drift per
5-minute window, worst-window versus first-window p99, and (Java) GC count,
GC pause total, and allocation rate over the run. A "zero-GC" claim requires a
soak row showing zero collections, not a headline row.

### `prod_parity`

Same as `headline` but with validation, persistent store, and dictionary
enabled for every engine, per each engine's documented production
configuration. This row answers "what do I actually get in production" and
guards against engines whose speed exists only in the everything-off
configuration. Never mixed into `headline` tables.

## Audit Pack

Every report-grade run should retain enough metadata to reproduce the row:

- scenario name and run profile
- transport: plain TCP or Onload (and Onload stack version/options)
- engine and version
- Java/C++ runtime and compiler flags
- command line and system properties
- warmup/measured iterations and runs
- throughput target and achieved throughput
- scheduler-delay distribution (see Load Model)
- final extraction-sink value (see Extraction Sink)
- admin messages processed during the measured window, per direction
- wire shape (`fixbench.wire.shape`)
- MD depth (`md.levels` and `ttt.md.levels`), and `stream.rate`,
  `stream.burst`, `stream.mark_every` for streaming rows
- core pinning strategy, CPU model, turbo/SMT state, and core isolation setup
- NIC model and IRQ affinity for cross-machine rows
- the matched `baseline_rtt` row id used for engine-cost decomposition
- Java GC/allocation counters for `soak` rows
- sample FIX messages for each direction, captured before measurement starts
- raw latency samples or HdrHistogram payload for CDF plots

If any implementation cannot satisfy the field-extraction contract for a
scenario, the row should be marked unsupported instead of being reported as a
latency comparison.
