Benchmarks and parity#
This section is the numerical record for quantem.gpu. It keeps current
results easy to find while preserving the detailed history needed to understand
which optimizations were accepted, rejected, or remain provisional.
Where to find the numbers#
Evidence |
Use it for |
|---|---|
Package contract and routes for API, kernel, and application developers; intentionally contains no benchmark tables |
|
The single friendly current overview of platform, bin, dtype, cache state, memory, parity, device, and test date |
|
Filterable measured, partial, pending, refuted, and unsupported configurations plus stable commands for closing each open gate |
|
CUDA 6 GiB VRAM and WebGPU 8 GB laptop acceptance status; runtime details on CUDA and WebGPU |
|
Metrics with date, revision, device, shape/dtype, cache state, load plan, benchmark definition, calibration, and parity provenance |
|
Latest documentation changes, implementation baselines, and why a retained number changed or remained historical |
|
Capability and implementation-source status without duplicated timing tables |
|
SwiftPM product structure, exact resident-summary contract, ownership boundary, and links to retained timing |
|
Direct-Metal output lifetime, resident-versus-driver memory accounting, profiling fields, and focused checks |
|
Required timing stages, cold/warm definitions, memory reporting, and acceptance rules |
|
Experimental count ANS versus bitpacking: GPU timings, calculated resident layouts, measured shared scratch, and unresolved per-codec/full-load peaks |
|
PR smoke, weekly physical profiles, manual signoff, comparison keys, run registry, and regression decisions |
|
Exact integer contracts, floating metrics, fixtures, and hardware gates |
|
Accepted and rejected IO, kernel, display, and browser experiments |
|
Real-data load/decode/product gates and backend signoff |
|
Shape-by-backend SSB kernel history, memory, and exactness |
|
Frozen historical extraction record; see native-size verification, experiment 20260915-metal-ssb-native-sizes (raw evidence in the private evidence archive), for the current 128/256/512 scan support and parity checks |
|
Pass-graph, redundant-work failure, and prevention checklist |
|
Physical low-memory Apple load/decode profiling and retained kernel evidence |
|
Archived browser G(q,k) layout and memory experiments |
The machine-readable fingerprints for retained numerical pages are available as
evidence_manifest.json. The manifest makes
an evidence edit explicit: changing a number requires updating its fingerprint
and rerunning the documentation guard.
The machine-readable execution schedule lives in
benchmarks/profile_matrix.json,
and completed, failed, refuted, or superseded runs remain indexed in the
RUNS.md registry of the private evidence archive.
Exact configuration coverage and agent runbooks live in
benchmarks/benchmark_registry.json.
Regenerate the filterable page with
python scripts/benchmark_registry.py render; CI fails if the generated table
drifts from the registry.
Reading a result#
A complete row answers all of these questions:
Which exact source and source revision ran?
What were the scan/detector shape, dtype, crop, bin, mask, and precision?
Which backend, device, driver/runtime, and kernel revision ran?
Was storage cold, controlled uncached-source-page, page-cache warm, process warm, or a saved-result reopen, and which audits/indexes already existed?
What did file open, read, decode, reduction, upload, synchronization, first usable product, and total wall time cost?
What were peak process memory, accelerator allocation/reserve, total-device occupancy, and swap/pressure?
Which frozen output or reference proves scientific parity?
If one of these facts is missing, the number may still be useful diagnostically but is not a release or migration signoff.
Current versus historical evidence#
The dashboard summarizes the current retained result; the results page owns its complete provenance. Maintainer ledgers keep chronological experiments—including regressions—so future work does not repeat failed kernel layouts or confuse an older record with the production path.
Dated pages with headings such as “Question,” “TODO,” or “Next” are collected under Historical experiments. Their wording is preserved as experiment provenance and does not define current commitments.
Valid older numbers are not copied through current overview pages. They remain once in the owning historical ledger with exact revision and protocol context. Measurements that failed parity or repeatability remain named as rejected experiments, never as current timing rows.
See later SSB performance checkpoints for subsequent full-aperture measurements. The original evidence remains frozen.