Testing and evidence#
Run the smallest gate that can disprove the change first, then widen coverage.
Portable Python#
python -m pip install -e ".[dev,movie]"
npm ci
PYTHONPATH=src python -m pytest -q # interactive tier, about a minute on one GPU
PYTHONPATH=src python -m pytest -q -m "" # everything, including the slow SSB fits (about 12 minutes)
Plain pytest skips tests marked slow (SSB fits and long hardware checks,
each several seconds or more). Mark a new test @pytest.mark.slow when it
takes more than about 3 s. CI and every release run the full set.
npm ci installs the pinned JavaScript development tools from
package-lock.json (esbuild, TypeScript, jsfive, WebGPU types) into
node_modules/. The WebGPU contract tests bundle the TypeScript sources with
them and run the bundles in Node.js 22 or newer; they are skipped only when
Node.js is not installed.
Portable CI proves imports, public contracts, CPU references, source packaging, and tests that do not require a physical accelerator. Backend skips are reported; they are not parity evidence.
Validate the profiling schedule and retained run registry in the same portable environment:
python scripts/check_profile_registry.py
python scripts/benchmark_registry.py validate
This checks that every parity capability has exactly one cell for CPU reference, CUDA, Python MPS, native Swift/Metal, and WebGPU; unsupported cells stay fail-closed; and terminal experiment manifests retain full revisions, input and output hashes, timestamps, and a human registry row. It also checks every exact benchmark gate, evidence import, runbook command, generated coverage table, and measured-row timing/parity requirement.
Choosing a reproducible performance gate#
Do not begin from a command copied out of an old experiment. Ask the registry for the open rows on the backend or computer you own:
python scripts/benchmark_registry.py next --platform "Python MPS"
python scripts/benchmark_registry.py next --computer "MacBook Air (M2, 8 GB)"
python scripts/benchmark_registry.py show io.mps.apple-m5-max-128gb.bin2.cold-original
python scripts/benchmark_registry.py command io.mps.apple-m5-max-128gb.bin2.cold-original
The command view prints the preflight, required environment, repository-owned
entry point, required artifacts, and promotion boundary. If it says
parity preflight, the command cannot produce a physical timing claim. If a
performance harness is missing, the coverage row stays partial or pending until
one is added.
The standard reusable entry points are:
Operation |
Entry point |
Primary output |
|---|---|---|
Python CUDA/MPS/CPU HDF5 load |
|
Run-level p50/p95/max, full-volume hash/dtype/shape parity, logical resident bytes, sampled accelerator/process peak, and release snapshots |
CUDA/MPS screening |
|
Separate build and prepared-reopen distributions, product hashes, stages, cache bytes, and memory |
Native Swift/Metal indexed load |
|
Exact resident volume/products, stage timing, hashes, and Metal/process memory |
Hardware-WebGPU HDF5 load |
|
Browser-local profile, checksums, run timing, and browser telemetry |
Python MPS SSB |
|
Reconstruction, parity-pair, and complete-fit JSONL |
Native Swift/Metal SSB |
|
Reconstruction, phase/loss, calibration, and cache-policy report |
Each physical run still needs a manifest and RUNS.md row in the private
evidence archive before launch; raw logs stay there, never in this repository,
and the dated summary (question, setup, numbers, conclusion) goes under docs/. The harness JSON is evidence, not permission to promote a dashboard
number without review.
Native Swift and Metal#
xcrun swift-format lint --strict --recursive \
Package.swift \
native/swift/Sources \
native/swift/Tests \
native/swift/Benchmarks
swift test
swift test -c release --filter MetalSSBKernelsTests
Opt-in real-source tests retain their fixture hashes and environment variables. A simulator or compile-only result does not replace physical Metal execution. The package-owned SSB benchmark accepts an exact metadata file, plane-major BF source, independent phase reference, repetition count, cache policy, fit-trial count, and fit-repetition count:
swift run -c release metal-ssb-benchmark \
METADATA_JSON FULL_BF_U8 REFERENCE_PHASE_F32 7 full 200 3
Record preparation separately from warm reconstruction, exact loss, and the complete fit. A finite cache budget is a different resource-policy row, not a replacement for the complete-cache timing.
Documentation#
python -m pip install -r docs/requirements.txt
jupyter-book build docs
The documentation build is hardware-independent and never downloads private data. It publishes retained evidence generated by qualified hardware runs.
Backend parity#
Follow Cross-backend parity. Re-run each unexpected hardware mismatch twice on the same source/environment before deciding whether it is deterministic drift or nondeterminism.
Never edit a frozen golden, expected hash, reference value, or tolerance merely to make a new implementation pass.
Hardware timing cadence and promotion rules are defined in Continuous profiling. A PR smoke test never becomes a physical-device performance claim.