# Cell-wide OPEN schedule v4 — implementation contract

Status: opt-in implementation and separate publication path available; installed
verification and the full sweep are in progress. Existing v3 observations retain
their original aligned, integer-floor schedule. Fresh v4 evidence lives in a
different TSV/page and never replaces a v3 observation implicitly.

## Configuration

Keep `--open-loop`. The launcher consumes this complete additional block and
exports the corresponding fixed environment variables to the client only:

| Flag | Environment variable |
| --- | --- |
| `--open-schedule cell-v4` | `LIBHFT_BENCH_OPEN_SCHEDULE` |
| `--cell-throughput R` | `LIBHFT_BENCH_OPEN_CELL_THROUGHPUT` |
| `--cell-sessions N` | `LIBHFT_BENCH_OPEN_CELL_SESSIONS` |
| `--client-index p` | `LIBHFT_BENCH_OPEN_CLIENT_INDEX` |
| `--client-processes P` | `LIBHFT_BENCH_OPEN_CLIENT_PROCESSES` |

Require Linux, OPEN nos_er, one measuring thread per process, `1 <= R <= 10^9`,
`1 <= P <= N <= 65535`, `N % P == 0`, `0 <= p < P`, local sessions `L=N/P`,
and existing numeric barrier ID `p+1`. Reject partial/contradictory configuration.
The legacy integer `--throughput` is a compatibility field equal to `floor(R/N)`;
v4 scheduling and reported target use the authoritative rational `R/N`, never
that floor. Initially require `R >= N`. Without any v4 fields, v3 is unchanged.

## Exact schedule and lateness

For process-owned ordinal `u` starting at zero, `q=p+P*u`, local session `u % L`,
and globally unique session `(u % L)*P+p`:

```
due_ns = epoch_ns + floor((q+1)*1_000_000_000/R)
phase_end_ns = epoch_ns + floor((N*iterations)*1_000_000_000/R)
```

Use checked quotient/remainder arithmetic, not repeated truncated intervals.
All times and products must fit signed 64-bit nanoseconds. At 50k/32/4 the cell,
process and session intervals are 20 us, 80 us and 640 us. The exact per-session
target is 1562.5/s; 20,000 iterations/session occupy 12.8 seconds.

Policy name: `bounded-catch-up`. Preserve the intended schedule and charge all
lateness to scheduler/response samples. Each loop performs a bounded receive
quantum (at most one nonblocking step per local session), then considers at most
one earliest process-owned slot. A step's already-published replies may be drained
with an explicit finite queue bound (.NET Native: at most 255), without another
transport step; otherwise deferred polling could manufacture a reply-ring overflow.
No per-session drain-of-debt loops, no reanchor,
and no new arbitrary lateness shedding threshold. A late generator can still
produce catch-up sends: actual task-start spacing must be measured and disclosed,
not claimed uniform merely from the intended formula. Existing in-flight cap
shedding, deficits, watchdog, correlation guards and saturation gates stay intact.

## Prepared phase barrier

Both `warmup` and `measured` rendezvous on every 1-based run. Allocate/touch sample
buffers, reset phase counters and validate empty request rings BEFORE readiness.
Warmup must exercise the service, response and scheduler sample stores too, in
discarded scratch buffers. Scratch-buffer presence is not a measured-phase flag:
warmup must not enter measured elapsed time, overload/peak counters, percentile
aggregation, extraction-sink publication or raw export. Reset or replace storage
before each phase's readiness without resetting live correlation identities.
Use exclusive ready-file creation; never follow/truncate an existing marker.
Warmup may have zero iterations (still rendezvous); measurement requires at least
one. Publish complete ready payloads in one write or atomically without replacing
an existing marker. The coordinator treats a zero-length ready file as pending.

```
ready.v4.<run>.<phase>.<barrier-id>
go.v4.<run>.<phase>
abort.v4
```

Ready and release payloads are ASCII newline-delimited `key=value`, with exactly
these shared keys, no duplicate/unknown/missing keys:

```
version=4
run=1
phase=warmup
clock=linux-monotonic-ns
policy=bounded-catch-up
cell_throughput=50000
cell_sessions=32
client_processes=4
iterations=20000
boot_id=<contents of /proc/sys/kernel/random/boot_id>
time_namespace=<readlink /proc/self/ns/time>
```

Ready additionally contains `client_index=p`. Release instead contains
`epoch_ns=<future CLOCK_MONOTONIC timestamp>`. Require exact agreement with the
prepared configuration, current boot and time namespace. Emit one
`bench_open_phase` diagnostic line with release metadata and client index after
validation. The harness validates exact participant IDs and payloads, then
atomically publishes each release with a two-second lead. Reject a release
already late at consumption; never fall back to a process-local origin. Timeout,
malformed payload or missing/exited peer creates `abort.v4`, never a success file.
Clients check abort while waiting. Preserve release/ready evidence in the run output.

C++ uses a benchmark-only CLOCK_MONOTONIC clock throughout the v4 pass, including
completion and watchdog timestamps; engine and v3 clocks remain unchanged. Java
and .NET use their verified scan Linux runtime monotonic clocks; .NET additionally
requires Stopwatch frequency 1,000,000,000. This is a pinned-runtime support claim,
not a portable cross-process guarantee of System.nanoTime or Stopwatch.

## Reporting and rollout

Time measured phases from their shared epoch through the later of phase end or
last completion; exclude barrier wait. Report per-session `target_per_sec` as a
decimal (1562.5 at 50k/32), with appended `schedule_version=4 target_num=50000
target_den=32 cell_target_per_sec=50000 schedule_policy=bounded-catch-up` fields.
For a non-terminating rational, round half-up to nine decimal places using integer
quotient/remainder arithmetic, then trim trailing zeroes and the decimal point.
The integer numerator/denominator remains authoritative.
Aggregate achieved rate keeps the existing MIN-per-session semantics. Preserve
failed partial/raw evidence; do not convert aborted or measured-shed phases into
headline successes.

Warmup retains v3 accounting: consuming all intended slots and draining every
actual send permits measurement even if warmup shed at the cap. Disclose warmup
shedding separately; do not infer zero warmup shedding from measured counters.
The no-shedding headline gate applies to measured work, not a new warmup SLO.

The coordinator, launcher, parameter auditor, raw/pacing replay, ingestion and
HTML provenance must recognize v4 explicitly before any v4 result replaces a v3
observation. Golden fixtures cover 50k/32/4 interleaving, 50k/1/1, rate-3 rounding,
phase boundaries, overflow, malformed/late/aborted releases and failed raw results.

## Run and publish the separate v4 experiment

Build and install every selected runtime first, including its native library.
On scan, keep orchestration on housekeeping cores; the machine configuration pins
the actual clients and servers to their declared measurement cores.

```sh
# Read-only selection, then a new full-size sitting (the output must not exist).
python3 bench/load/run-open-matrix.py --host scan --schedule-version 4 --list
taskset -c 0-9 python3 bench/load/run-open-matrix.py --host scan \
  --schedule-version 4 --out /home/yann/libhft-bench-runs/open-v4-full-NEW
```

Use `--only '^java-pure/kernel/single$'` for a scoped full-size run. Resuming an
interrupted sitting requires the same version, selection, recipe and installed
artifacts; use a fresh directory for a rebuild or a retry of a completed cell.
The driver rechecks source and installed-artifact hashes before every launch and
before publishing its result; a mid-sitting change interrupts instead of silently
mixing binaries. Its importer receives the same scrubbed environment as its audits.
The driver updates `open-sweep-scan-v4.json`, the separate v4 matrix and the
handbook after every completed, audited cell. Version 3 remains the CLI default.

To publish an existing completed **full-size** capture explicitly:

```sh
python3 bench/load/ingest-rows.py /path/to/cell/rows.csv scan \
  --mode open --schedule-version 4
python3 bench/load/gen-matrix.py scan --schedule-version 4
python3 bench/load/gen-handbook.py
```

This writes `docs/bench/results-scan-open-v4.tsv` and
`docs/bench/latency-matrix-scan-open-v4.html` and the current main page
`latency-matrix-scan.html`. The original `*-open.tsv` remains v3, rendered separately
as `latency-matrix-scan-legacy-v3.html`. Numeric v4 publication requires the standard
timed capture -- 30 s measured and 5 s warmup per run, three runs, each session offered
its exact share of the 50,000/s cell aggregate, so a run measures 1,500,000 messages at
any session count (standards 4, 2026-09-12; rows captured earlier used 20,000 warmup /
20,000 measured iterations × three runs, and every row records the window it measured) --
exact schedule/phase provenance, actual affinity read-back, and complete raw
sample replay. Smoke and nonstandard recovery probes do not enter this table.
All eight displayed service/response statistics are reconstructed from complete
raw series and checked against the per-session run means and aggregate maxima;
unchecked two-decimal harness CSV values are not the publication authority.
The importer recognizes one cold `bench_percentiles method=nearest-rank-integer-v1`
stamp per client for exact integer nearest-rank selection. Unstamped legacy
captures retain their emitter's original arithmetic; duplicate, malformed, late
or mixed-process method stamps are rejected. No tolerance silently accepts either
adjacent p99.9 rank. The correction has passed the language tests and protocol audit
and is installed for the new full-size sweep; historical functional captures
predate it.
Full-delivery saturation may retain visibly qualified diagnostic values; it is
not a passing capacity claim. Maximum throughput remains a separate rate sweep.
