Benchmarks and parity#

This section is the numerical record for quantem.gpu. It keeps current results easy to find while preserving the detailed history needed to understand which optimizations were accepted, rejected, or remain provisional.

Where to find the numbers#

Evidence

Use it for

Documentation landing page

Package contract and routes for API, kernel, and application developers; intentionally contains no benchmark tables

Implementation dashboard

The single friendly current overview of platform, bin, dtype, cache state, memory, parity, device, and test date

Coverage and runbooks

Filterable measured, partial, pending, refuted, and unsupported configurations plus stable commands for closing each open gate

Minimum-device gates

CUDA 6 GiB VRAM and WebGPU 8 GB laptop acceptance status; runtime details on CUDA and WebGPU

Verified benchmark results

Metrics with date, revision, device, shape/dtype, cache state, load plan, benchmark definition, calibration, and parity provenance

Revision and change ledger

Latest documentation changes, implementation baselines, and why a retained number changed or remained historical

Backend coverage

Capability and implementation-source status without duplicated timing tables

Native Swift and Metal

SwiftPM product structure, exact resident-summary contract, ownership boundary, and links to retained timing

Python MPS

Direct-Metal output lifetime, resident-versus-driver memory accounting, profiling fields, and focused checks

Benchmark methodology

Required timing stages, cold/warm definitions, memory reporting, and acceptance rules

CUDA codec speed and peak memory

Experimental count ANS versus bitpacking: GPU timings, calculated resident layouts, measured shared scratch, and unresolved per-codec/full-load peaks

Continuous profiling

PR smoke, weekly physical profiles, manual signoff, comparison keys, run registry, and regression decisions

Cross-backend parity

Exact integer contracts, floating metrics, fixtures, and hardware gates

Optimization ledger

Accepted and rejected IO, kernel, display, and browser experiments

Load acceptance evidence

Real-data load/decode/product gates and backend signoff

SSB performance evidence

Shape-by-backend SSB kernel history, memory, and exactness

Native Swift/Metal SSB migration

Frozen historical extraction record; see native-size verification, experiment 20260915-metal-ssb-native-sizes (raw evidence in the private evidence archive), for the current 128/256/512 scan support and parity checks

Native Metal loader postmortem

Pass-graph, redundant-work failure, and prevention checklist

M2 Air Metal evidence

Physical low-memory Apple load/decode profiling and retained kernel evidence

WebGPU memory history

Archived browser G(q,k) layout and memory experiments

The machine-readable fingerprints for retained numerical pages are available as evidence_manifest.json. The manifest makes an evidence edit explicit: changing a number requires updating its fingerprint and rerunning the documentation guard.

The machine-readable execution schedule lives in benchmarks/profile_matrix.json, and completed, failed, refuted, or superseded runs remain indexed in the RUNS.md registry of the private evidence archive.

Exact configuration coverage and agent runbooks live in benchmarks/benchmark_registry.json. Regenerate the filterable page with python scripts/benchmark_registry.py render; CI fails if the generated table drifts from the registry.

Reading a result#

A complete row answers all of these questions:

  1. Which exact source and source revision ran?

  2. What were the scan/detector shape, dtype, crop, bin, mask, and precision?

  3. Which backend, device, driver/runtime, and kernel revision ran?

  4. Was storage cold, controlled uncached-source-page, page-cache warm, process warm, or a saved-result reopen, and which audits/indexes already existed?

  5. What did file open, read, decode, reduction, upload, synchronization, first usable product, and total wall time cost?

  6. What were peak process memory, accelerator allocation/reserve, total-device occupancy, and swap/pressure?

  7. Which frozen output or reference proves scientific parity?

If one of these facts is missing, the number may still be useful diagnostically but is not a release or migration signoff.

Current versus historical evidence#

The dashboard summarizes the current retained result; the results page owns its complete provenance. Maintainer ledgers keep chronological experiments—including regressions—so future work does not repeat failed kernel layouts or confuse an older record with the production path.

Dated pages with headings such as “Question,” “TODO,” or “Next” are collected under Historical experiments. Their wording is preserved as experiment provenance and does not define current commitments.

Valid older numbers are not copied through current overview pages. They remain once in the owning historical ledger with exact revision and protocol context. Measurements that failed parity or repeatability remain named as rejected experiments, never as current timing rows.

See later SSB performance checkpoints for subsequent full-aperture measurements. The original evidence remains frozen.