Revision and change ledger#
This ledger separates documentation revisions from the implementation revision that produced a benchmark. A documentation edit never makes an older timing a measurement of the current code.
2026-09-08: unpublished native packed-resident candidate#
The dated investigation records exact original-HDF5 reconstruction, seven-resident detector reductions, native presentation measurements, rejected experiments, and corrected timing assumptions. It is not a release signoff or a replacement for the older source checkpoints below. The 120 Hz wide-detector and four-to-five-second seven-load targets remain unmet.
2026-08-22: previous review#
Ledger reviewed: 2026-08-22
Integration base: origin/main
75be74eCurrent clean benchmark source checkpoints: canonical Python/CPU metadata and parity
68dbe3a; canonical native Swift/Metal load5106ca4Historical measured checkouts: consolidated Python MPS
0bc9378; native controlled-sourcec0ea444; retained CUDA and WebGPU rows remain bound to the exact revisions shown in the results ledgerDocumentation branch:
benchmark-platform-computer-matrix
The native stack extends exact prepared QH5 binning (e0e92b4), optimized
word-major detector binning (ff3c7fd), and exact resident summaries
(d65911a) with bounded streaming, private residency, and controlled macOS
source-page measurement through c0ea444. It does not relabel an older CUDA,
Python MPS, WebGPU, source-load, or application measurement.
Latest documentation changes#
Commit |
Change |
Performance-number effect |
|---|---|---|
|
Normalized source and working detector metadata, then reran canonical CPU/Python MPS bins 1/2/4/8 |
replaces the |
|
Added the canonical cross-layout logical-pixel contract used by native Swift/Metal evidence |
adds current physical M5 24 GB bin2/bin4 distributions and one pressure-gated full-native bin1 smoke; the bin1 smoke is partial because it caused 758,448,128 B of swapouts |
|
Consolidated exact MPS bin1/2/4/8 kernels, overflow and partial-tail guards, resident reuse, and exception-safe cleanup |
adds final-head post-warmup fresh-destination p50 0.414824/0.457153/0.382109/0.356258 s for bins 1/2/4/8; controlled cold and application E2E remain open |
|
Adopted the accepted exact WebGPU integer CoM/DPC/iDPC source without importing obsolete campaign harnesses |
no new timing claim; preserves the existing physical-Apple numerical evidence at the consolidated source boundary |
|
Added scratch-free exact full- |
historical pre-consolidation MPS rows remain retained; they no longer headline current source |
Current MPS lifecycle-audit follow-up |
Separates resident payload, driver allocation sampled after load and release, process RSS, pressure, and swap |
replaces the pressure-contaminated 2026-08-19 MPS load headline with revision |
|
Aligned the raw benchmark state with explicit macOS |
promotes a controlled uncached-source-page full-native package row; its distribution replaces schema v7 only because the older raw state was contradictory, not because the IO path changed |
|
Reused the immutable source identity while preserving controlled source reads |
kept the same exact source/load boundary; the later |
|
Added exact full-native detector-bin-1 loading into private Metal residency |
established the 18 GiB exact resident path; later controlled-source runs own the current distribution |
|
Added provenance-bound exact resident summaries and overflow-safe |
adds separate one-time summary-build and prepared-reopen rows; does not replace first compressed-source load or resident-compute rows |
|
Added |
adds separate native Swift/Metal SSB rows; does not replace differently configured CUDA, Python MPS, or WebGPU rows |
Current profiling-registry follow-up |
Added the 35-cell platform/module schedule, human run index, manifest validator, CI gate, and continuous-profiling guide |
none; the existing timings and scientific acceptance states are unchanged |
|
Combined current platform profiling evidence with the three parity-qualified scientific fixes |
existing measured baselines remain tied to their recorded revisions |
|
Clarified that the SSB PyTorch functions are executable teaching references, not production kernels |
none |
|
Rewrote SSB examples with explicit named layouts and ordinary PyTorch FFT calls |
none |
|
Made the SSB reference progressive and executable from top to bottom |
none |
|
Added equation-paired PyTorch SSB reference functions |
none |
|
Added the implementation and benchmark dashboard |
no measurement changed; existing evidence was indexed |
|
Completed operation contracts, source maps, and parity gates |
none |
|
Turned API pages into typed integration contracts |
none |
|
Separated kernel implementations from remote compute |
none |
|
Standardized public |
none |
|
Reorganized the site around scientific operations and runtimes |
none |
Local documentation commits are recorded as branch-relative identifiers until they are published. Do not construct public GitHub commit links for an unpublished revision.
Latest benchmark-evidence changes#
This table records promotion decisions, not another copy of the measurements. The current values and distributions live only in Verified benchmark results.
Evidence change |
Current disposition |
Why |
|---|---|---|
|
Promoted for Python MPS exact warm-process and independent-process loading at its revision; |
Four-bin full-output parity is byte exact. ABBA measurements preserve process/source state and explicit release. Bins 2/4/8 meet strict p50 <=0.5 s; full native bin1 is 0.523 s. Source pages remain warm and uncontrolled, so the result is not cold storage or application E2E. |
|
Promoted for Python MPS process-state and memory reporting |
Seven fresh processes per detector bin and seven explicit-release warm-process repetitions preserve exact selected-frame hashes, record logical payload plus driver allocation sampled after load and release, process RSS, pressure, and swap, and pass independent exact detector-bin/product parity. Source pages were warm and uncontrolled, so neither protocol is cold-storage evidence. |
|
Superseded, preserved |
Its repeated-load harness cleared the Torch cache but did not call |
|
Promoted as controlled native package evidence |
Seven fresh processes use a new index root, fresh private destination, and explicit |
|
Superseded, preserved |
Its command applied |
|
Promoted as separate native Swift/Metal prepared-product evidence |
Seven fresh-process reopens reproduce nine same-device products byte-for-byte on MacBook Pro (M5 Max, 128 GB) and MacBook Air (M2, 8 GB). Summary creation, prepared reopen, resident load, compressed-source load, and GUI paint remain distinct boundaries. |
|
Promoted as a separate native Swift/Metal SSB row |
It uses its own full-BF fixture and optimizer implementation; warm prepared compute is not raw-source load, application wall time, or physical 8 GB signoff. |
|
Current cross-platform profile |
It supplies atomic timing, memory, dtype/bin, device, date, and parity fields; fixtures C and D remain explicitly separate. |
Exact streamed screening follow-up |
Promoted |
It derives masks from the complete detector sum and transparently reruns BF/DF when the provisional mask differs. The first-chunk-mask candidate failed full-scan parity and remains rejected. |
Deterministic CUDA SSB calibration follow-up |
Promoted |
Three seeded fits reproduce fitted parameters, phase, object, and loss. The earlier atomic-objective run produced two fitted minima and remains rejected despite being faster. |
Current MPS prepared companion |
Rejected |
Stored columns do not match the declared detector-bin coordinate grid; only the raw detector-bin-2 reconstruction retains current numerical evidence. |
Current WebGPU profile |
Promoted for the measured load and resident-compute boundaries |
UI paint, physical 8 GB signoff, per-pixel CoM/DPC/iDPC error arrays, and calibration remain explicit gaps. |
|
Superseded for current headline timing |
It remains useful same-fixture history, but newer comparable profile rows own current values. |
Detector-bin-4 products |
Mixed |
Integer products and CoM pass their stated gates; cross-backend iDPC remains blocked and cannot inherit an integer-parity check mark. |
Native-detector MPS CoM/iDPC |
Blocked as native-resolution evidence |
The public interaction sidecar uses detector bin 2; direct full-resolution CoM passes, while iDPC remains blocked. |
|
Separate application evidence |
It proves a physical 8 GB detector-bin-4 application path, not a no-bin library load or native SSB memory gate. |
|
Historical, labeled first-process |
The storage-cache eviction procedure was not retained, so it cannot be called cold. |
|
Historical single visible run |
No timing distribution was retained, so it is not a median. |
Other July rows with missing host identity |
Historical diagnostic |
Their numerical evidence can guide optimization, but incomplete hardware provenance prevents current release signoff. |
Accepted and rejected performance experiments remain in the optimization ledger; older SSB layout experiments remain in the SSB performance history.
Adding a new revision#
Add the complete benchmark row with date, exact measured revision, physical device/runtime, source shape/dtype, cache state, crop/bin/load plan, timing boundary, memory, calibration, and parity artifact.
Explain whether it replaces a truly comparable row or remains separate.
Update the dashboard only after the detailed row is complete.
Update the evidence manifest when a fingerprinted evidence page changes.
Keep rejected or incomparable experiments discoverable rather than deleting them.