# 32-session pacing review — scan, 2026-09-08

The mono, pool and shard clients divide the target correctly but do not generate
an evenly spaced aggregate stream. All eight sessions within each process start
with the same deadline; each process chooses its own post-barrier origin. Overdue
slots catch up back-to-back. These are properties of the existing v3 workload,
not proof of a transport bottleneck.

## Retained-capture replay

The offline checker [pacing_audit.py](../../bench/load/pacing_audit.py) reconstructs
local task starts from chronological scheduler samples under the **explicit**
aligned-v3 schedule assumption:

`relative_t0[j] = (j + 1) * 640204 + scheduler_delay_ns[j]`

For client process 1, run 1, the first 1,000 slots of each of eight sessions, all
15 retained kernel cells have short within-batch gaps and long between-batch gaps.
Each entry below is the nearest-rank **median / p90 adjacent local task-start gap
in microseconds**. Uniform spacing within one 12,496/s client process would be
80.0255 microseconds.

| Flavour | Mono | Pool | Shard |
| --- | ---: | ---: | ---: |
| C++ | 4.999 / 597.241 | 5.044 / 600.879 | 5.224 / 597.400 |
| Java Pure | 6.546 / 586.514 | 6.521 / 587.785 | 6.589 / 586.266 |
| Java/JNI | 6.369 / 588.195 | 6.739 / 588.547 | 6.704 / 588.351 |
| .NET Native | 6.411 / 593.303 | 6.895 / 585.492 | 6.442 / 587.500 |
| .NET Pure | 5.977 / 593.852 | 6.085 / 588.035 | 5.878 / 594.540 |

[pacing-review-scan.json](pacing-review-scan.json) records the exact input paths,
SHA-256 hashes, checker hashes, raw-count validation, assumptions and statistics.
It uses the R10 replacement for native kernel/shard, not its withdrawn R7 result.
The three managed captures have since been withdrawn from canonical latency
publication for missing CPU-pin enforcement. They remain diagnostic evidence of
the send schedule's shape here, not isolated-core performance measurements.
The complete replay is retained at
`/home/yann/libhft-bench-runs/pacing-review-20260908/kernel-15-cells.json`.

The checker requires complete per-session/run scheduler, response and service
series, zero recorded drops and shedding, full replies, and unclamped scheduler
values. It checks that reconstructed starts are nondecreasing over the entire
selected run, even when reporting only a prefix. Raw benchmark names must match
the client engine. Java explicitly reports zero stale replies; absent C++/.NET
stale counters remain **unreported**, not zero.

These conditions do not prove source provenance, input pairing or writer order:
those remain explicit caller assumptions, requiring the retained source review.
The archived C++/.NET writers append scheduler samples on the send path and save
them before sorting latency copies; receive-side uncertainty does not reorder
that send-side series. Counts alone would not prove this.

This is **not** packet capture: `t0` precedes encoding and sending. There is no
captured common epoch across processes, so the replay does not infer whole-cell
spacing. Selected-prefix distributions are not whole-run distributions or new
latency results. No verdict threshold is imposed and no matrix row is imported.

Example, using the retained Java shard capture:

```sh
taskset -c 0-9 python3 bench/load/pacing_audit.py \
  --recipe /home/yann/libhft-bench-runs/open-v3-20260908-r7/shard/java-pure/kernel/attempt-07i2f284/run/recipe.txt \
  --client-log /home/yann/libhft-bench-runs/open-v3-20260908-r7/shard/java-pure/kernel/attempt-07i2f284/run/java-pure_kernel_1.log \
  --raw /home/yann/libhft-bench-runs/open-v3-20260908-r7/shard/java-pure/kernel/attempt-07i2f284/run/raw/shard_java-pure_kernel_1.hftbrw \
  --interval-ns 640204 --assume-schedule aligned-per-session-v3 --run 1 --slots 1000
```

## Shared-phase requirements, not an implemented policy

A staggered schedule needs more than different local session offsets. The
current barrier writes an empty `go.<run>` file; clients poll its presence every
millisecond. It establishes readiness, not a common clock epoch.

C++ `FixUtils::GetTimeNs()` selects `CLOCK_MONOTONIC_RAW` on this build platform.
Disassembly of the installed Java 17 and .NET 8 runtimes confirms that their
timers use `CLOCK_MONOTONIC`; an inert .NET probe reports exactly 1,000,000,000
Stopwatch ticks/second. Nanosecond units do not make RAW and MONOTONIC timestamps
interchangeable. The pinned-runtime proof, hashes, commands and probe source are
retained in `pacing-review-20260908/clock-review/README.md` under the run directory.
Java's portable API is not a cross-JVM clock guarantee.

For a future, versioned scheme, each homogeneous cell needs a declared common
clock domain, a future per-run epoch, unique global session indices, and an exact
phase/rounding formula. Publish the structured epoch atomically; distinguish
successful release from abort. Prepare buffers and resets before readiness, or
provide enough lead time and explicitly fail/disclose late arrivals. Never
silently substitute a fresh local origin. Preserve the requested/effective rate
distinction: v3 currently rounds 50,000/32 to 1,562, or 49,984/s effective.

The phase-policy choice remains pending. No runtime schedule or old result has
been silently relabelled by this review.

## Separate .NET reply-integrity correction

Source review also found that both .NET open loops indexed timestamps using only
`replyId & 2047`, without validating full-ID slot ownership. A small outstanding
count alone cannot prevent wraparound over one old unanswered request. Managed
handling could retain a previous ID when tag 11 was absent; both clients had overly
permissive numeric-ID parsing. MD intermediate responses also needed validation
before folding or sending derived orders.

Both open clients now use the shared `OpenRequestRing`: reserve the complete ID
and timestamps before transport, reject live-slot aliasing, validate the exact
wire ID and expected reply stage before extraction, and consume terminal replies
once. Generations remain continuous across warmup and measured passes. A native
initial-send refusal cancels only its own reservation; a refused derived order
fails the already-started exchange. Managed duplicate FIX sequence numbers use
the same open-only failure path. Correlation failures retain valid partial
samples, stop later phases and emit failed raw/result footers.

Verification: 64 shared-library tests and 101 multi-client tests passed, including
the actual socket-free managed decoder and native decode/fold helper. The new
`dotnet_open_audit.py` checks production wiring, with mutation tests covering
guards moved out of scope or after extraction/sending. Protocol parity passes
147 checks. This does not substitute for a fresh installed-artifact benchmark;
the retained matrix rows above predate the correction.

The first installed kernel smoke, retained under
`dotnet-open-correlation-smoke-20260908-Ena5b5`, delivered all replies in both
flavours' single/mono/pool/shard cells: 3,000 warmup and 3,000 measured requests
per session, no deficits, shedding or correlation failures. It did **not** pass
the whole smoke audit: native single and managed pool exceeded the scheduler-slip
gate. More importantly, managed multi/open bypasses the shared runner's CPU-pin
application, despite accepting `--measure-core`; all four managed client pin
checks lack evidence. The smoke layout checker also carries the preceding
single-session expected client set into native pool. These placement defects are
fixed. Retained-source manifests confirm that all twelve managed OPEN rows share
the affinity omission; their latency/rate values are now withdrawn pending fresh
pinned runs, while original records remain in `withdrawn_results`. See
[DOTNET-OPEN-AFFINITY-REVIEW.md](DOTNET-OPEN-AFFINITY-REVIEW.md). Native single's server footer
also reports a terminal-state error after client completion. Review found that
the worker continued calling `Step` after its first negative return and overwrote
that initial result with lifecycle error -4. Both worker paths now stop immediately
and preserve the first code/status. They do not relabel the initial -1 as success:
native `Step` can use it for either EOF or a protocol/read error. This correction
has now had installed kernel smoke verification: the follow-up sitting
`/home/yann/libhft-bench-runs/dotnet-affinity-smoke-20260908-p1vR5e` completed all
eight .NET kernel cells with startup affinity readback passing. Seven passed the
smoke gate; native single remained scheduler-saturated (p99 delay 199.842 us).
All 582,000 requested measured replies arrived, with zero measured shedding or
deficits. The native single footer preserved the first `-1` / `Disconnected`
instead of overwriting it with `-4`; this still does not classify the initial
cause as harmless. These one-run kernel probes do not replace full-size,
all-transport measurements. No smoke latencies have been imported.

This fixes a source-level risk, **not evidence that bad replies occurred in the
retained captures**. Existing successful delivery counts do not prove the missing
guards were exercised. Rate, pacing/catch-up policy, outstanding cap and the
two-second watchdog are unchanged; the separate phase-policy choice remains pending.
