Testing and evidence#

Run the smallest gate that can disprove the change first, then widen coverage.

Portable Python#

python -m pip install -e ".[dev,movie]"
npm ci
PYTHONPATH=src python -m pytest -q          # interactive tier, about a minute on one GPU
PYTHONPATH=src python -m pytest -q -m ""    # everything, including the slow SSB fits (about 12 minutes)

Plain pytest skips tests marked slow (SSB fits and long hardware checks, each several seconds or more). Mark a new test @pytest.mark.slow when it takes more than about 3 s. CI and every release run the full set.

npm ci installs the pinned JavaScript development tools from package-lock.json (esbuild, TypeScript, jsfive, WebGPU types) into node_modules/. The WebGPU contract tests bundle the TypeScript sources with them and run the bundles in Node.js 22 or newer; they are skipped only when Node.js is not installed.

Portable CI proves imports, public contracts, CPU references, source packaging, and tests that do not require a physical accelerator. Backend skips are reported; they are not parity evidence.

Validate the profiling schedule and retained run registry in the same portable environment:

python scripts/check_profile_registry.py
python scripts/benchmark_registry.py validate

This checks that every parity capability has exactly one cell for CPU reference, CUDA, Python MPS, native Swift/Metal, and WebGPU; unsupported cells stay fail-closed; and terminal experiment manifests retain full revisions, input and output hashes, timestamps, and a human registry row. It also checks every exact benchmark gate, evidence import, runbook command, generated coverage table, and measured-row timing/parity requirement.

Choosing a reproducible performance gate#

Do not begin from a command copied out of an old experiment. Ask the registry for the open rows on the backend or computer you own:

python scripts/benchmark_registry.py next --platform "Python MPS"
python scripts/benchmark_registry.py next --computer "MacBook Air (M2, 8 GB)"
python scripts/benchmark_registry.py show io.mps.apple-m5-max-128gb.bin2.cold-original
python scripts/benchmark_registry.py command io.mps.apple-m5-max-128gb.bin2.cold-original

The command view prints the preflight, required environment, repository-owned entry point, required artifacts, and promotion boundary. If it says parity preflight, the command cannot produce a physical timing claim. If a performance harness is missing, the coverage row stays partial or pending until one is added.

The standard reusable entry points are:

Operation

Entry point

Primary output

Python CUDA/MPS/CPU HDF5 load

scripts/benchmark_hdf5_load.py

Run-level p50/p95/max, full-volume hash/dtype/shape parity, logical resident bytes, sampled accelerator/process peak, and release snapshots

CUDA/MPS screening

scripts/benchmark_screening.py

Separate build and prepared-reopen distributions, product hashes, stages, cache bytes, and memory

Native Swift/Metal indexed load

metal-4dstem-indexed-load-benchmark

Exact resident volume/products, stage timing, hashes, and Metal/process memory

Hardware-WebGPU HDF5 load

scripts/benchmark_webgpu_h5_browser.py

Browser-local profile, checksums, run timing, and browser telemetry

Python MPS SSB

scripts/benchmark_ssb_mps_scaling.py

Reconstruction, parity-pair, and complete-fit JSONL

Native Swift/Metal SSB

metal-ssb-benchmark

Reconstruction, phase/loss, calibration, and cache-policy report

Each physical run still needs a manifest and RUNS.md row in the private evidence archive before launch; raw logs stay there, never in this repository, and the dated summary (question, setup, numbers, conclusion) goes under docs/. The harness JSON is evidence, not permission to promote a dashboard number without review.

Native Swift and Metal#

xcrun swift-format lint --strict --recursive \
  Package.swift \
  native/swift/Sources \
  native/swift/Tests \
  native/swift/Benchmarks
swift test
swift test -c release --filter MetalSSBKernelsTests

Opt-in real-source tests retain their fixture hashes and environment variables. A simulator or compile-only result does not replace physical Metal execution. The package-owned SSB benchmark accepts an exact metadata file, plane-major BF source, independent phase reference, repetition count, cache policy, fit-trial count, and fit-repetition count:

swift run -c release metal-ssb-benchmark \
  METADATA_JSON FULL_BF_U8 REFERENCE_PHASE_F32 7 full 200 3

Record preparation separately from warm reconstruction, exact loss, and the complete fit. A finite cache budget is a different resource-policy row, not a replacement for the complete-cache timing.

Documentation#

python -m pip install -r docs/requirements.txt
jupyter-book build docs

The documentation build is hardware-independent and never downloads private data. It publishes retained evidence generated by qualified hardware runs.

Backend parity#

Follow Cross-backend parity. Re-run each unexpected hardware mismatch twice on the same source/environment before deciding whether it is deterministic drift or nondeterminism.

Never edit a frozen golden, expected hash, reference value, or tolerance merely to make a new implementation pass.

Hardware timing cadence and promotion rules are defined in Continuous profiling. A PR smoke test never becomes a physical-device performance claim.