# hp CPU hardware epoch: Xeon Gold 6254

Recorded 2026-09-18. hp's processor was replaced. This record draws a hard
provenance boundary: it reports the initial native-tuner candidate below, not
an end-to-end performance result, and it does not reinterpret an earlier one.

## Decision

All hp floor, smoke and load-benchmark evidence captured before this change is
the **Xeon Gold 6154 epoch**. It remains useful historical evidence, but it
must not be blended with, overwritten by, or numerically compared to a Xeon
Gold 6254 run. A new hp result begins only after a fresh, separately labelled
6254 sitting has passed its preflight.

The live machine declaration, [`load-bench-machine-hp.conf`](../../bench/load/load-bench-machine-hp.conf),
now names the new hardware. Its changed fingerprint prevents a new recipe from
claiming the old machine declaration by accident. Historical result directories
and records must retain their original files and fingerprints.

## Before and after

| Property | Historical hp epoch | Current hp epoch, observed 2026-09-18 |
| --- | --- | --- |
| Processor | Intel Xeon Gold 6154 @ 3.00 GHz | Intel Xeon Gold 6254 @ 3.10 GHz |
| Socket / physical cores / SMT | 1 / 18 / off | 1 / 18 / off |
| Online CPUs / NUMA | historical runs used CPUs 0–17 / one node | CPUs 0–17 / node 0 contains CPUs 0–17 |
| Shared L3 reported by `lscpu` | do not infer from the replacement | 24.8 MiB |
| CPU identity | retained by each historical recipe and record | family 6, model 85, stepping 7; microcode `0x5003901` |

The unchanged core count and single-NUMA layout mean the placement plan itself
does not need to be redesigned. That is not a performance equivalence claim:
cache, frequency and processor-generation differences still require every
latency and floor number to be recaptured.

## Current benchmark-machine preflight

Observed directly after the replacement:

| Item | Current state |
| --- | --- |
| Kernel | `5.14.0-687.44.1.el9_8.x86_64` |
| Isolated CPUs | `8-17`; `nohz_full=8-17`; `rcu_nocbs=8-17`; IRQ affinity `0-7` |
| Pin declaration | coordinator 8; server workers 9–12; VMA internal thread 13; client cores 14–17; housekeeping 0–7 |
| CPU policy | `intel_pstate`, `performance`, 1.2–3.1 GHz on every policy |
| Boost state | `intel_pstate/no_turbo=1`; the kernel reports turbo disabled by BIOS or unavailable on this processor |
| Solarflare path | SFC9220 at `0000:2d:00.0/.1`; `sfcA`/`sfcB` (`ens5f0`/`ens5f1`) are UP and the DAC physical-link verification passed |
| Mellanox path | ConnectX-4 Lx at `0000:15:00.0/.1`; `mlxA`/`mlxB` (`ens4f0np0`/`ens4f1np1`) are UP and the DAC physical-link verification passed |

The effective boot isolation is:

```text
isolcpus=domain,managed_irq,8-17 nohz_full=8-17 rcu_nocbs=8-17 irqaffinity=0-7
```

No firmware or frequency-policy setting was changed while recording this
inventory. In particular, a fresh baseline must record the observed boost state
rather than treating a disabled boost as a benchmark defect or enabling it
mid-campaign.

## Initial C++ native-tuner candidate: rejected by the real-FIX gate

The Phase-0 C++ tuner completed on the new processor after its full
kernel-correctness corpus and decision-rule self-tests passed. It used GCC
14.2.1 from `gcc-toolset-14`, which resolves `-march=native` and `-mtune=native`
to `cascadelake`, including AVX-512 VNNI. The live measurement selected isolated
core 8 at 100% idle; it recorded the current governor and disabled-turbo policy.

The retained raw sample bundle and a compact result table are described in
[`evidence/hp-cpu-upgrade-20260918/TUNER.md`](evidence/hp-cpu-upgrade-20260918/TUNER.md).
The selected candidates are:

| Probe | Selected candidate | Measured change against the previous default |
| --- | --- | --- |
| Tag parse | unrolled | 9.915 vs 10.920 cyc/op, 9.2% faster |
| Single-byte scan | AVX2 | 4.947 vs 7.542 cyc/op, 34.4% faster |
| Integer format | SSE2 clang-asm form | 23.945 vs 26.803 cyc/op, 10.7% faster |
| Decimal format | clean arms | 19.488 vs 20.813 cyc/op, 6.4% faster |
| Structural parser | guarded AVX2 index | 1,585 vs 1,763 cyc/op for the 10-level-MD index/materialize comparison, 10.1% faster |

Checksum remains `vpaddb_masktail`: despite the new processor's AVX-512 VNNI,
its duty-cycle comparison measured 219.012 ns/message for that default versus
219.223 for the AVX-512 candidate. The copy, timestamp, integer-parse,
decimal-parse and field-split defaults also remained selected.

The complete real-FIX in-context gate subsequently **rejected** this candidate:
the full profile had a +2.73% median component regression and 83 material
regressions across three interleaved runs. The structural-parser enablement was
the clearest non-transfer: its standalone component profile increased
`parse_md_incremental` by 94.8%. The profile-free default build remains the hp
baseline; no staged binary or round-trip comparison was created. The detailed
gate record is in the linked tuner evidence.

The structural probe is operationally expensive (about 15 minutes of the
approximately 20-minute primitive run); schedule a complete rerun rather than
interpreting a partial structural trace.

## What is historical now

- `floor-*-hp.tsv`, `results-hp*.tsv`, `open-sweep-hp*.json`, and the hp
  latency-matrix HTML pages document the 6154 epoch. The hp matrix prose now
  labels that boundary without changing its data.
- The 2026-09-18 Onload pool-stall investigation was run on the 6154. Its
  mechanism and diagnostic method remain relevant, but neither its tail sizes
  nor its `EF_BUZZ_USEC` experiment establish 6254 performance.
- No previously published hp value is withdrawn solely because the hardware was
  replaced. It is archived with its actual machine, rather than presented as a
  value from the new machine.

## Required refresh order

1. Recheck this inventory after each reboot: CPU identity, microcode, isolation,
   `nohz_full`, `rcu_nocbs`, IRQ placement, governor/frequency limits and both
   network-rig physical links.
2. Build and install the selected current revision through the benchmark staging
   workflow. Do not run a current checkout beside root-staged artifacts from an
   earlier sitting.
3. Run the normal configuration, parity and live-pinning smoke preflights. They
   establish runnable placement and delivery; they are not latency evidence.
4. Re-run the remaining hp primitive, layout, codec, FFI and wire-floor suites.
   Store the resulting raw data and summaries under a distinct 6254 epoch name;
   do not overwrite the 6154 files.
5. Screen any future native-tuner candidate through the in-context gate and
   then the real `bench_roundtrip` guard before staging it. The first 6254
   candidate failed the former, so it must not be retested as a staged profile
   unless its prescription or gate outcome materially changes.
6. Capture a complete new load sitting with the current timed schedule: three
   runs, 5 s warmup and 30 s measured per cell. Publish only cells satisfying
   the normal delivery, pinning and variation rules.
7. Before treating the Onload pool tail as unchanged, repeat its baseline →
   candidate → baseline control on the intended flavour, load and topology, and
   retain raw maxima for warmup as well as measured phases.

The general benchmark contracts still apply; this record adds the hardware
identity boundary that must be satisfied before hp can again be used as a
current comparison host.
