Benchmark coverage and runbooks#

This page is the operational index for QuantEM.GPU performance work. It keeps required configurations visible even when they have never run, imports every retained atomic measurement from the current evidence index, and maps each open gate to a repository-owned runbook.

The registry separates three concepts:

  • a coverage gate defines the exact module, platform, computer, geometry, dtype, cache state, timing boundary, parity, and memory evidence required;

  • a retained measurement records what actually ran, including p50, p95, maximum, resident bytes, measured peaks, swap, device, date, and revision;

  • a runbook gives a stable command, preflight, required environment, and artifacts for repeating or closing one gate.

Status is not support by implication

A pending row is planned work. A partial row retains useful evidence but does not satisfy the full gate. A portable parity command never becomes physical timing evidence, and a prepared reopen never becomes a cold-source claim.

Agent entry point#

Start by validating the registry, then ask it for the next open gate. Do not guess a command from an old experiment page.

python scripts/benchmark_registry.py validate
python scripts/benchmark_registry.py next --limit 10
python scripts/benchmark_registry.py next --computer "MacBook Air (M2, 8 GB)"
python scripts/benchmark_registry.py next --platform "Python MPS"
python scripts/benchmark_registry.py show io.mps.apple-m5-max-128gb.bin2.cold-original
python scripts/benchmark_registry.py command io.mps.apple-m5-max-128gb.bin2.cold-original

Use --performance-entrypoint-only to hide gates whose current repository entry point is parity-only:

python scripts/benchmark_registry.py next --performance-entrypoint-only

Before a physical run, create the experiment manifest and RUNS.md row in the private evidence archive, as described in Continuous profiling. The command output intentionally does not launch hardware on its own: device ownership, fixture identity, cold-source control, and output locations must be resolved first.

Promotion rule#

A measured row needs all of the following under one comparison key:

  1. the exact source and implementation revisions;

  2. computer, accelerator, runtime, and device ownership;

  3. source and selected scan geometry, detector geometry, bin, crop, and dtype;

  4. explicit cold, warm, prepared, resident, or saved-result state;

  5. run-level records with p50, p95, maximum, and the wall-clock boundary;

  6. logical resident bytes, accelerator allocation/peak, process or browser-tree peak, total-device peak when available, pressure, and swap;

  7. independent scientific parity and output fingerprints; and

  8. a retained manifest, raw artifact, and terminal RUNS.md status.

If a field is unavailable, the row remains partial or pending. Calculated payloads are useful planning data but never replace measured peaks.

Current qualified load measurements#

Only explicitly designated current rows appear here. Historical, superseded, refuted, and unmeasured configurations remain in the complete tables below rather than being silently deleted.

Platform

Computer

State

Selected scan

Source detector

Detector bin

Output detector

Source dtype

Staging dtype

Resident dtype

Scientific gate

Cache/process state

Wall boundary

Samples

p50

p95

Maximum

Logical resident

Accelerator/driver peak

Process/tree peak

Process physical-footprint peak

Swap delta

Parity

Device tested

Date tested

Revision

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

one same-process warmup; operating-system source pages uncontrolled; no eviction performed; fresh returned destination released after each trial

public io.load return after MPS synchronization; full-volume hash and release excluded

7

0.406624 s

0.428164 s

0.428164 s

18.000 GiB

18.442 GiB

18.692 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

fresh QH5 index root and destination; macOS F_NOCACHE on source hashing and indexed descriptors; immutable source already audited

catalog, pipeline compilation, plan, complete exact private-resident volume, seven products, metadata, and provenance

7

0.577793 s

0.900979 s

0.900979 s

18.000 GiB

18.571 GiB

0.874 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

c0ea44465e6346a8436a0b74f491a04af0b5dc32

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.225984 s

0.231192 s

0.231192 s

4.500 GiB

5.087 GiB

0.657 GiB

5.336 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

6

0.200036 s

0.204694 s

0.204694 s

1.125 GiB

1.712 GiB

1.039 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

exact

original compressed source reread every visit; existing QH5 index and source-bound packing layout; exact DPC sums reused after first load; OS pages uncontrolled

indexed source open through synchronous complete packed-resident return; catalog and independent full-count audit separate

7

1.496660 s

2.086119 s

2.086119 s

18.000 GiB

n/a

n/a

2.531 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-09-06

e305f9216ed359397e27f69882931ffe16de8d99

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.678703 s

0.697179 s

0.697179 s

4.500 GiB

5.087 GiB

0.662 GiB

5.271 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.640942 s

0.651931 s

0.651931 s

1.125 GiB

1.712 GiB

1.018 GiB

2.319 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared immutable block indexes; explicitly warm source pages; fresh browser target per retained run; 699 pageouts; zero swap growth; no dropped runs

navigation through scientifically usable exact resident output and diagnostic frame checksums; exhaustive full-volume hash and application E2E excluded

7

1.358000 s

1.594000 s

1.594000 s

18.000 GiB

n/a

6.500 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared indexes; source pages unspecified; CPU exact-reference validation only

public CPU reference load; full-volume hash excluded

1

32.788039 s

32.788039 s

32.788039 s

18.000 GiB

n/a

19.141 GiB

n/a

0 B

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

Dtype and residency coverage#

Each row is one atomic source, staging, and resident dtype contract. Exact rows preserve scientific counts; browse-only rows are explicit saturating representations and cannot satisfy an exact gate.

Platform

Computer

State

Selected scan

Source detector

Scan bin

Detector bin

Output detector

Crop

Source dtype

Staging dtype

Resident dtype

Scientific gate

Precision contract

Implementation basis

Next gate or reason

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/cuda/kernels/bslz4.cu shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=cuda) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained.

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public CUDA io.load and retain its complete-volume hash, source/resident provenance and peak memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The CUDA decoder has a dedicated uint16 bitshuffle path and existing unit parity; the current registry has no uncontended full-volume native uint16 distribution.

Run the public CUDA load entry point on an uncontended device and retain a complete-volume uint16 hash plus source/staging/resident provenance and peak allocation.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form.

CUDA io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output.

CUDA io.load has no uint8 output dtype.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output.

CUDA io.load has no saturating uint8 browse output.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening.

CUDA io.load has no uint16-to-uint32 widening.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for CUDA.

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/mps/kernels/bslz4.msl shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=mps) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained.

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public Python MPS io.load and retain its complete-volume hash, source/resident provenance and peak memory.

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The MPS uint16 decoder and retained full-volume canonical hash preserve native counts at detector bin 1.

n/a

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form.

Python MPS io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output.

Python MPS io.load has no uint8 output dtype.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output.

Python MPS io.load has no saturating uint8 browse output.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening.

Python MPS io.load has no uint16-to-uint32 widening.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for Python MPS.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

Native4DSTEMIO can describe one-byte indexed sources, but Metal4DSTEMIndexedLoader.swift requires a uint16 source and uint16 staging at the public integrated boundary.

The integrated indexed Swift/Metal loader does not admit a native uint8 source.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

Metal4DSTEMIndexedLoader.swift supports the uint16 fallback stage, while retained optimized measurements used separately registered audited uint8 staging.

Run a source whose complete audit does not authorize uint8 staging and retain exact uint16 staging-to-resident parity and allocation evidence.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

NativeHDF5Bridge.swift describes one- and two-byte indexed sources, and Metal4DSTEMIndexedLoader.swift requires uint16 input.

The native indexed Swift/Metal source contract does not admit uint32 HDF5 detector values.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

Metal4DSTEMExactBinner.provenance accepts only uint16 or uint32 output, although it can use an audited uint8 staging buffer.

The exact native Swift/Metal load contract cannot publish a uint8 resident scientific volume.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

The indexed exact loader exposes uint16 or uint32 scientific output and has no separate saturating browse-resident API.

Explicit saturating uint8 resident output is not implemented by the public native Swift/Metal load boundary.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

Metal4DSTEMExactBinner has tested uint16-to-uint32 kernels and provenance, but Metal4DSTEMIndexedBinnedLoad fixes integrated resident output to uint16.

Expose uint32 output through the public indexed loader and cache contract, then retain complete load, metadata, and physical parity.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

Native4DSTEMValueRangeAudit binds source identity and maximum counts; the retained controlled full-native run used audited uint8 staging and reproduced exact uint16 volume and product hashes.

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

h5reader.ts, bslz4.ts, and local-h5.ts recognize matching native uint8 source, decode, and resident modes; no retained full-volume physical uint8 fixture proves the complete path.

Run an exact physical hardware-browser full-volume uint8 source through local HDF5 decode and retain output hash, allocation, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The physical WebGPU path retained exact native uint16 resident checks and seven runs, but device allocation and timed complete-volume hashing remain incomplete.

Add WebGPU device-allocation telemetry and retain an exact complete-volume hash inside the timed boundary before promoting the integrated load gate.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

h5reader.ts, bslz4.ts, and local-h5.ts implement matching native uint32 source/decode/resident modes; physical full-volume proof is absent.

Run a native uint32 HDF5 source on physical hardware WebGPU and retain the complete resident hash, browser/device memory, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

bslz4.ts has an audit-dependent low8 kernel, but local-h5.ts does not bind a complete audit identity to the returned uint8 resident state.

Replace the experimental low8 global flag with a typed source-identity-bound audit in the local-HDF5 public result, then retain exact full-volume parity.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

bslz4.ts decodes all uint16 planes and saturates to 255; source tests distinguish this from the experimental low8 audit path.

Run the fused full-plane clip8 path on physical hardware with values above 255 and retain output hash, browser/device memory, and browse-only provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

local-h5.ts requires decode dtype uint32 to match a uint32 HDF5 source and rejects a uint16 source request.

The WebGPU local-HDF5 public path does not widen uint16 source values into uint32 resident storage.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

local-h5.ts keeps decode and resident integer mode matched for native paths and rejects mismatched uint32 requests; it exposes no uint8-stage-to-uint16 reconstruction contract.

Audited uint8 staging followed by exact uint16 resident reconstruction is not implemented in the public WebGPU local-HDF5 path.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/cpu.py returns the HDF5 native dtype unchanged at detector bin 1; a complete public-API uint8 fixture artifact is not retained.

Add a repository fixture that proves public io.load preserves every native uint8 count, shape, order, and provenance field at detector bin 1.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The CPU reference returns native uint16 counts at detector bin 1 and has retained physical reference probes; the portable registry gate is not a complete public-API artifact.

Promote only after the portable public API fixture retains complete-volume uint16 hash, shape, order, mask, and metadata parity.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

The CPU reference preserves the HDF5 native dtype at detector bin 1; public load defaults may auto narrow uint32, so the exact uint32 request must be explicit and tested.

Add a public-API native uint32 fixture that disables advisory auto narrowing and proves complete count, dtype, shape, and metadata parity.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint8 output.

CPU reference io.load has no uint8 output dtype.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no saturating uint8 browse output.

CPU reference io.load has no saturating uint8 browse output.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint32 widening.

CPU reference io.load has no uint16-to-uint32 widening.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for CPU reference.

Detector-product coverage#

Each row fixes one platform, reproducible computer, detector bin, and exact product-suite boundary. Partial diagnostics retain their sample and memory context, but only fully measured gates display p50, p95, and maximum timing.

Platform

Computer

State

Selected scan

Source detector

Detector bin

Output detector

Resident dtype

Cache/process state

Wall boundary

Samples

p50

p95

Maximum

Logical resident

Accelerator/driver peak

Process/tree peak

Swap delta

Parity

Device tested

Date tested

Revision

Next gate or reason

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

warm operating-system source pages; unspecified; empty screening-result cache

quantem.gpu.screening.prepare package wall through exact complete six-array screening result in balanced A-B-B-A-B-A-A-B comparison

6

n/a

n/a

n/a

n/a

2.084 GiB

1.958 GiB

0 B

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

023a6c497b106b216c87205d3fbec63377d77177

Run the current public ScreeningResult on an uncontended CUDA device with benchmark_screening.py –require-exact-full-suite; retain all nine hashes, uint64 total/ABF/ADF dtype and shape, p50/p95/maximum, and complete memory telemetry.

Python MPS

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before decoder, allocator, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming product path that never requires an 18 GiB resident input before attempting this physical row.

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

exact resident detector-bin-1 MPS input; source pages uncontrolled or warm; no prepared product result; five pageouts; zero swapout growth

single first-execution diagnostic only; synchronized detector-product publication was 0.499909 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

27.510 GiB

20.436 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b8df61b55920cb098e2848301f6d30c45870511d

Repeat the exact product suite without pageout growth for synchronized p50, p95, and maximum timing plus complete memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native MPS input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run only after an uncontended pressure preflight; use a single exact full-native smoke with a hard stop on swap/pageout growth before considering repeated product timing.

Native Swift/Metal

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB private-resident input exceeds the computer’s 8 GB unified memory before decoder, Metal, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming or mapped product path that never requires an 18 GiB private-resident input before attempting this physical row.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 Metal input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the complete product suite from the explicit exact detector-bin-2 resident volume with full parity, synchronized timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 Metal input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the complete exact detector-bin-4 product suite on the current clean revision with full hashes, synchronized timing, pressure, compressed memory, swap, and peak memory; detector bin 2 remains the locked application acceptance policy.

Native Swift/Metal

MacBook Air (M2, 8 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

prepared immutable QH5 indexes; source pages unspecified

complete exact resident volume, products, metadata, calibration, and provenance; full private readback/hash excluded

6

0.313870 s

0.318865 s

0.318865 s

18.000 GiB

18.571 GiB

0.639 GiB

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

531e5001f5a0a886c058f772ef42d770f000890b

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.225984 s

0.231192 s

0.231192 s

4.500 GiB

5.087 GiB

0.657 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

6

0.200036 s

0.204694 s

0.204694 s

1.125 GiB

1.712 GiB

1.039 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified

package exact indexed load plus exact full-scan products; app UI, index creation, and full-volume hash excluded

1

n/a

n/a

n/a

18.000 GiB

18.571 GiB

0.624 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

0b305f69fc1039933463986a0f74609e22d4dd35

Repeat uncontended full-native publication with complete product hashes, p50/p95/max, driver allocation, RSS, pressure, and swap.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.678703 s

0.697179 s

0.697179 s

4.500 GiB

5.087 GiB

0.662 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.640942 s

0.651931 s

0.651931 s

1.125 GiB

1.712 GiB

1.018 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB total memory before browser, adapter, decoder, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming browser product path that never requires an 18 GiB resident input before attempting this physical row.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

prepared QH5 indexes; warm source pages; exact full-native uint16 resident input; fresh browser target per retained run

complete internal GPU product suite from exact resident input through awaited readbacks; parity hashing, sorting, and encoding excluded

7

0.482000 s

0.483600 s

0.483600 s

18.000 GiB

n/a

6.492 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native hardware-WebGPU input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run a guarded physical hardware-browser admission smoke, then retain the complete native product suite with full parity, timing, browser/device memory, pressure, and swap if admission remains safe.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native CPU input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the independent full-native CPU reference product suite only when host CPU, storage, and memory ownership are uncontended; retain all hashes, timing, and peak RSS.

CPU reference

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before products and operating-system headroom. Next: Implement and prove an exact bounded-streaming CPU reference product path that never requires an 18 GiB resident input before attempting this physical row.

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

exact full-native CPU-resident input after a warm operating-system page-cache source load

single independent detector-product traversal was 31.078040 seconds; source load excluded; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

36.450 GiB

n/a

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-19

334b7b5135fe29787540370a00f280fa138430a2

Add ABF, retain all product hashes, and repeat the exact physical CPU reference traversal for p50/p95/maximum and complete process-memory telemetry.

CPU reference

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native CPU input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run only after an uncontended memory preflight; retain one exact full-native CPU reference smoke, complete product hashes, pressure, swap, and peak RSS before repetition.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

resident reference input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Keep exact reference tests current; no CPU speed claim is required.

Coverage summary#

State

Gate count

✓ Measured

18

◐ Partial

34

○ Pending

73

! Blocked

35

× Refuted

4

Not supported

27

↺ Superseded

0

Platform

Tracked gates

CUDA

17

Python MPS

37

Native Swift/Metal

56

WebGPU

58

CPU reference

23

Platform and computer coverage#

Each row identifies one reproducible hardware configuration. Counts describe tracked cells, including explicit unsupported contracts; a pending value remains a test to run. Load, admission, memory, and performance gates are multiplied across compatible computers because hardware changes the result. Platform-wide correctness or unsupported contracts are recorded once instead of creating misleading duplicate hardware rows.

Platform

Computer

Tracked cells

Measured

Partial

Pending

Blocked

Refuted

Unsupported

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

17

2

5

5

0

0

5

Python MPS

MacBook Air (M2, 8 GB)

9

0

0

7

2

0

0

Python MPS

MacBook Pro (M5 Max, 128 GB)

18

2

3

6

0

2

5

Python MPS

MacBook Pro (M5, 24 GB)

10

0

1

8

0

1

0

Native Swift/Metal

MacBook Air (M2, 8 GB)

14

0

0

11

2

0

1

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

27

8

5

7

0

0

7

Native Swift/Metal

MacBook Pro (M5, 24 GB)

15

4

3

7

0

0

1

WebGPU

MacBook Air (M2, 8 GB)

15

0

0

4

11

0

0

WebGPU

MacBook Pro (M5 Max, 128 GB)

27

2

4

7

9

1

4

WebGPU

MacBook Pro (M5, 24 GB)

16

0

0

7

9

0

0

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

2

0

0

2

0

0

0

CPU reference

MacBook Air (M2, 8 GB)

2

0

0

0

2

0

0

CPU reference

MacBook Pro (M5 Max, 128 GB)

2

0

2

0

0

0

0

CPU reference

MacBook Pro (M5, 24 GB)

2

0

0

2

0

0

0

CPU reference

Portable CI runner

15

0

11

0

0

0

4

Required coverage gates#

Every row is one exact scientific and device configuration. A pending row is work to do, not an implicit failure and not evidence that a backend is supported on that device.

Platform

Computer

State

Module

Operation

Selected scan

Source detector

Detector bin

Output detector

Source dtype

Staging dtype

Resident dtype

Scientific gate

Cache/process state

Wall boundary

Priority

Runbook

Next gate

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

CoM, DPC, and iDPC

CoM row/column, rotation, and iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident CUDA inputs

complete synchronized product publication

2

dpc-parity

Add a stable physical timing entry point and retain complete full-precision parity, synchronized timing, and memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm streamed screening source

complete exact product suite

2

detector-product-profile

Run the current public ScreeningResult on an uncontended CUDA device with benchmark_screening.py –require-exact-full-suite; retain all nine hashes, uint64 total/ABF/ADF dtype and shape, p50/p95/maximum, and complete memory telemetry.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident CUDA maps

synchronized display-ready outputs

3

display-fft-parity

Add a stable CUDA display-kernel harness for first execution and warm publication with parity and memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident CUDA maps

synchronized Fourier output

3

display-fft-parity

Add a stable CUDA FFT harness for first execution and warm publication with parity and memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

I/O and load

cold arbitrary-source full-native load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold arbitrary source; control not yet qualified

source open through exact resident output

1

python-load-matrix

Run controlled cold-source discovery/read/decode/load distributions without disturbing unrelated GPU owners.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

I/O and load

warm-source full-native resident load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm source; historical p50 exists but complete current distribution is missing

first usable exact resident output

2

python-load-matrix

Re-run the current clean revision for seven trials with p95, maximum, process RSS, total-card peak, and exact hashes.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

I/O and load

native uint16 exact residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

source and prepared state declared by the run

complete exact resident output

3

python-load-matrix

Run the public CUDA load entry point on an uncontended device and retain a complete-volume uint16 hash plus source/staging/resident provenance and peak allocation.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

I/O and load

native uint8 exact residency

512 × 512

192 × 192

1

192 × 192

uint8

uint8

uint8

exact

source and prepared state declared by the run

complete exact resident output

3

python-load-matrix

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public CUDA io.load and retain its complete-volume hash, source/resident provenance and peak memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

I/O and load

audited uint8 staging to exact uint16 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

Audited uint8 staging with uint16 resident reconstruction is not implemented for CUDA.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

I/O and load

native uint32 exact residency

512 × 512

192 × 192

1

192 × 192

uint32

uint32

uint32

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

CUDA io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

I/O and load

exact uint16 to uint32 widening

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint32

exact

source and prepared state declared by the run

complete exact widened resident output

4

registry-smoke

CUDA io.load has no uint16-to-uint32 widening.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

I/O and load

audited lossless uint16 to uint8 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint8

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

CUDA io.load has no uint8 output dtype.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

I/O and load

explicit saturating uint16 to uint8 browse residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint8

browse-only

source and prepared state declared by the run

complete saturating browse resident output

4

registry-smoke

CUDA io.load has no saturating uint8 browse output.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

one-pass exact screening products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm source; empty result cache

complete six-array screening result build

2

screening-profile

n/a

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

prepared immutable screening-result reopen

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable screening-result cache; fresh-process reopen

complete saved-result reopen; source build excluded

3

screening-profile

n/a

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

2

cuda-ssb-preflight

Add a stable physical calibration harness and retain complete fit timing, parameters, repeatability, parity, and memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

cuda-ssb-preflight

Add a stable CUDA performance entry point and retain physical full-size timing, memory, phase/loss parity, and hashes.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

CoM, DPC, and iDPC

CoM row/column, rotation, and iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS inputs

complete synchronized product publication

2

dpc-parity

Run the physical MPS entry point with complete full-precision parity, synchronized timing, pressure, swap, and peak memory.

Python MPS

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admissible

blocked exact-residency contract

4

detector-product-profile

Implement and prove an exact bounded-streaming product path that never requires an 18 GiB resident input before attempting this physical row.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS maps

synchronized display-ready outputs

3

display-fft-parity

Run first-execution and warm MPS display-kernel distributions with full parity, synchronized publication, pressure, swap, and peak memory.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS maps

synchronized Fourier output

3

display-fft-parity

Run first-execution and warm MPS FFT distributions with full parity, synchronized publication, pressure, swap, and peak memory.

Python MPS

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

full-native resident admission

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admitted on the 8 GB resident path

admission decision before allocation

1

python-load-matrix

Keep this fail-closed until a bounded exact streaming result can preserve the full logical tensor without materializing an 18 GiB resident destination.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

Screening

one-pass exact screening products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

declared source-page state; empty result cache

first usable and exact complete screening result build

1

screening-profile

Run bounded-staging MPS screening construction with full-source product parity, source-stage timing, pressure, swap, and p50/p95/max.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

Screening

prepared immutable screening-result reopen

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable screening-result cache

complete saved-result reopen; source build excluded

2

screening-profile

Run a fresh-process prepared-result reopen distribution with complete product hashes, cache bytes, RSS, compressed memory, swap, and p50/p95/max.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

2

mps-ssb

After the bounded object smoke passes, run the frozen fit with deterministic repeatability, fitted-parameter parity, pressure, swap, and memory.

Python MPS

MacBook Air (M2, 8 GB)

○ Pending

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

mps-ssb

Run one bounded smoke, then the frozen 512 workflow with full phase/loss parity, pressure, swap, memory, and repeat distributions.

Python MPS

MacBook Pro (M5 Max, 128 GB)

× Refuted

CoM, DPC, and iDPC

CoM row/column, rotation, and iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident full-native MPS inputs

complete synchronized product publication

1

dpc-parity

Diagnose the full-native iDPC disagreement without weakening the frozen 2e-5 maximum-error gate, then rerun the complete synchronized suite with p50/p95/maximum and memory telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact full-native MPS input

complete product suite publication

2

detector-product-profile

Repeat the exact product suite without pageout growth for synchronized p50, p95, and maximum timing plus complete memory telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS maps

synchronized display-ready outputs

3

display-fft-parity

Add a stable MPS display-kernel harness for first execution and warm publication with parity and memory.

Python MPS

MacBook Pro (M5 Max, 128 GB)

× Refuted

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident full-precision MPS map

synchronized Fourier output

1

display-fft-parity

Diagnose the Torch-MPS FFT disagreement against the frozen float32 reference without weakening tolerances, then rerun first-execution and warm FFT distributions with complete parity and memory telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

cold original compressed-source first encounter

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold original compressed source; control not yet qualified

source open through exact complete resident output

1

python-load-matrix

Run seven independently controlled cold-source trials with stage timing, complete parity, and full memory telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

native uint16 exact residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

source and prepared state declared by the run

complete exact resident output

3

python-load-matrix

n/a

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

native uint8 exact residency

512 × 512

192 × 192

1

192 × 192

uint8

uint8

uint8

exact

source and prepared state declared by the run

complete exact resident output

3

python-load-matrix

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public Python MPS io.load and retain its complete-volume hash, source/resident provenance and peak memory.

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

post-warmup exact package load with fresh returned destination

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

one same-process warmup; operating-system source pages uncontrolled; fresh returned destination released after each trial

public io.load return after MPS synchronization; full-volume hash and release excluded

3

python-load-matrix

n/a

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

exact warm package load with explicit destination reuse

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm or uncontrolled source pages; explicit resident destination recycled

public package load return after final Metal completion

3

python-load-matrix

Repeat the recycled-destination lifecycle with full-volume canonical-layout parity and complete release telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

audited uint8 staging to exact uint16 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

Audited uint8 staging with uint16 resident reconstruction is not implemented for Python MPS.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

native uint32 exact residency

512 × 512

192 × 192

1

192 × 192

uint32

uint32

uint32

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

Python MPS io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

exact uint16 to uint32 widening

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint32

exact

source and prepared state declared by the run

complete exact widened resident output

4

registry-smoke

Python MPS io.load has no uint16-to-uint32 widening.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

audited lossless uint16 to uint8 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint8

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

Python MPS io.load has no uint8 output dtype.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

explicit saturating uint16 to uint8 browse residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint8

browse-only

source and prepared state declared by the run

complete saturating browse resident output

4

registry-smoke

Python MPS io.load has no saturating uint8 browse output.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

Screening

one-pass exact screening products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

declared source-page state; empty result cache

first usable and exact complete screening result build

2

screening-profile

Run complete MPS screening construction with full product parity, source-stage timing, memory, and p50/p95/max.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

Screening

prepared immutable screening-result reopen

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable screening-result cache

complete saved-result reopen; source build excluded

3

screening-profile

Run a fresh-process prepared-result reopen distribution with complete product hashes, cache bytes, RSS, and p50/p95/max.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

1

mps-ssb

Run the frozen 200-trial plus Nelder-Mead workflow on physical MPS with deterministic repeatability and full memory telemetry.

Python MPS

MacBook Pro (M5 Max, 128 GB)

○ Pending

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

mps-ssb

Run native 512 evidence with the frozen reference engine, full parity, memory, and repeat distributions.

Python MPS

MacBook Pro (M5, 24 GB)

◐ Partial

CoM, DPC, and iDPC

CoM row/column, rotation, and iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS inputs

complete synchronized product publication

2

dpc-parity

Repeat the complete public DPC pipeline without display contention or pageout growth for synchronized p50, p95, maximum, and complete memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact full-native MPS input

complete product suite publication

1

detector-product-profile

Run only after an uncontended pressure preflight; use a single exact full-native smoke with a hard stop on swap/pageout growth before considering repeated product timing.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS maps

synchronized display-ready outputs

3

display-fft-parity

Run first-execution and warm transform, histogram, and color distributions with frozen parity and complete memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

× Refuted

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident MPS maps

synchronized Fourier output

1

display-fft-parity

Diagnose the Torch-MPS FFT disagreement against the frozen float32 reference without weakening tolerances, then rerun first-execution and warm FFT distributions with complete parity and memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

cold original compressed-source first encounter

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold original compressed source; no prepared result

source open through exact complete resident output

1

python-load-matrix

After the bounded admission smoke passes, run independently controlled cold-source trials with full-native parity, pressure, paging, and swap telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

post-warmup exact package load with fresh returned destination

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

one same-process warmup; source pages declared by run; fresh destination

public io.load return after backend synchronization

2

python-load-matrix

Run one bounded admission smoke before seven retained trials; sample driver allocation, process high-water, pressure, and swap without recycling the destination.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Screening

one-pass exact screening products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

declared source-page state; empty result cache

first usable and exact complete screening result build

2

screening-profile

Run bounded-staging MPS screening construction with full-source product parity, source-stage timing, memory, and p50/p95/max.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Screening

prepared immutable screening-result reopen

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable screening-result cache

complete saved-result reopen; source build excluded

3

screening-profile

Run a fresh-process prepared-result reopen distribution with complete product hashes, cache bytes, RSS, and p50/p95/max.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

1

mps-ssb

Run the frozen 200-trial plus Nelder-Mead workflow with deterministic repeatability, fitted-parameter parity, and full memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

mps-ssb

Run the frozen 512 workflow with full phase/loss parity, memory, and repeat distributions.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

CoM, DPC, and iDPC

rotation and Fourier iDPC from resident CoM maps

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident float32 CoM row and column maps

synchronized rotation and iDPC publication

2

dpc-parity

Run the current native Metal path with frozen full-precision parity, synchronized timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

resident exact detector-bin-2 Metal input

complete product suite publication

1

swift-indexed-load

Run the complete product suite from the explicit exact detector-bin-2 resident volume with full parity, synchronized timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

resident exact detector-bin-4 Metal input

complete product suite publication

2

swift-indexed-load

Run the complete exact detector-bin-4 product suite on the current clean revision with full hashes, synchronized timing, pressure, compressed memory, swap, and peak memory; detector bin 2 remains the locked application acceptance policy.

Native Swift/Metal

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admissible

blocked exact-residency contract

4

swift-indexed-load

Implement and prove an exact bounded-streaming or mapped product path that never requires an 18 GiB private-resident input before attempting this physical row.

Native Swift/Metal

MacBook Air (M2, 8 GB)

Not supported

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

not applicable

unsupported load-plan contract

4

swift-indexed-load

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized display-ready outputs

3

display-fft-parity

Run the current Metal display-kernel distribution with full parity, first-versus-warm timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized Fourier output

3

display-fft-parity

Run the current Metal FFT distribution with full parity, first-versus-warm timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

full-native resident admission

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admitted on the 8 GB resident path

admission decision before allocation

1

swift-indexed-load

Keep this fail-closed until an exact bounded residency contract preserves full logical coverage without allocating the 18 GiB destination.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

I/O and load

cold original compressed-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold original source; no prepared-result substitution

source open through exact complete 4.5 GiB working volume and products

1

swift-indexed-load

Run coordinated physical 8 GB trials with full scan, no crop, exact uint16 sums, pressure, RSS, compressed memory, and swap.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

I/O and load

prepared-index exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable index; source pages declared per run

exact complete 4.5 GiB working volume and products

1

swift-indexed-load

Run repeated uncontended physical 8 GB prepared-index loads and prove admission, parity, memory, and recovery.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

I/O and load

cold original compressed-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold original source; no prepared-result substitution

source open through exact complete working volume and products

2

swift-indexed-load

Run pressure-guarded physical cold-source trials with exact sums, full geometry/calibration provenance, RSS, compressed memory, and swap.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

I/O and load

prepared-index exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable index; source pages declared per run

exact complete working volume and products

2

swift-indexed-load

Run repeated uncontended physical prepared-index loads with exact parity, memory, cancellation, and recovery evidence.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

2

swift-ssb

After the bounded object smoke passes, run the full fit with fitted-parameter parity, deterministic repeatability, pressure, swap, memory, and timing.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

swift-ssb

Run one bounded smoke, then the current native 512 workflow with exact phase/loss parity, pressure, swap, memory, and repeat distributions.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

CoM, DPC, and iDPC

rotation and Fourier iDPC from resident CoM maps

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident float32 CoM row and column maps

synchronized rotation and iDPC publication

3

dpc-parity

Run the current native Metal path with frozen full-precision parity, synchronized timing, first-versus-warm states, and peak memory.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused

exact indexed load, detector sum, and seven native-resolution products

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused

exact indexed load, detector sum, and seven native-resolution products

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared exact resident input

complete exact resident volume and seven products

3

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

not applicable

unsupported load-plan contract

4

swift-indexed-load

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized display-ready outputs

3

display-fft-parity

Register the current Metal display-kernel distribution with exact revision, device, parity, peak memory, and first-versus-warm states.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized Fourier output

3

display-fft-parity

Register the current Metal FFT distribution with exact revision, device, parity, peak memory, and first-versus-warm states.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

audit-free arbitrary-source cold first encounter

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold arbitrary source; no audit or index prepared

source selection through exact complete resident output and products

1

swift-indexed-load

Measure discovery, value-range audit, index creation, load, products, and peak memory from an unaudited source without relabeling controlled F_NOCACHE evidence.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

cold arbitrary-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold arbitrary source; no audit or index prepared

source selection through exact complete working output and products

1

swift-indexed-load

Measure discovery, value-range audit, index creation, exact sum-bin-2 load, products, and peak memory from an unaudited source.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

controlled uncached exact private-resident load and products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

fresh QH5 index root and destination; F_NOCACHE; already-audited immutable source

catalog through exact complete resident volume and seven products

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

prepared-index exact private-resident load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

cold arbitrary-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold arbitrary source; no audit or index prepared

source selection through exact complete working output and products

2

swift-indexed-load

Measure discovery, value-range audit, index creation, exact sum-bin-4 load, products, and peak memory from an unaudited source.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

audited uint8 staging to exact uint16 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

3

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

native uint16 exact residency with uint16 staging

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

source and prepared state declared by the run

complete exact resident output

3

swift-indexed-load

Run a source whose complete audit does not authorize uint8 staging and retain exact uint16 staging-to-resident parity and allocation evidence.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

exact uint16 to uint32 widening

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint32

exact

source and prepared state declared by the run

complete exact widened resident output

3

registry-smoke

Expose uint32 output through the public indexed loader and cache contract, then retain complete load, metadata, and physical parity.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

prepared-index exact private-resident load and products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable QH5 indexes; source pages unspecified

complete exact resident volume, products, metadata, calibration, and provenance

3

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

prepared-index exact private-resident load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

3

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

native uint32 exact residency

512 × 512

192 × 192

1

192 × 192

uint32

uint32

uint32

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

The native indexed Swift/Metal source contract does not admit uint32 HDF5 detector values.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

native uint8 exact residency

512 × 512

192 × 192

1

192 × 192

uint8

uint8

uint8

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

The integrated indexed Swift/Metal loader does not admit a native uint8 source.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

audited lossless uint16 to uint8 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint8

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

The exact native Swift/Metal load contract cannot publish a uint8 resident scientific volume.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

explicit saturating uint16 to uint8 browse residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint8

browse-only

source and prepared state declared by the run

complete saturating browse resident output

4

registry-smoke

Explicit saturating uint8 resident output is not implemented by the public native Swift/Metal load boundary.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

exact detector-bin-8 resident load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

not applicable

unsupported contract test

5

registry-smoke

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

Screening

prepared screening workflow

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not applicable

unsupported contract test

5

registry-smoke

The backend parity contract currently marks the prepared screening workflow not implemented for native Swift/Metal.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

Selective loading

64 by 64 evidence-selective rectangle

64 × 64

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source-page state declared

exact selected resident output

2

swift-indexed-load

Implement and expose the backend-neutral bulk selection contract, then prove source-range selectivity, order, metadata, parity, bytes, memory, and timing.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

Selective loading

arbitrary evidence-selective positions with order and duplicates

n/a

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source-page state declared

exact ordered selected resident output

2

swift-indexed-load

Implement the canonical selector contract without UI policy, then prove evidence-selective IO, order, duplicates, parity, bytes, memory, and timing.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

2

swift-ssb

Register a current clean full-fit distribution with fitted-parameter parity, repeatability, cache policy, and memory.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

swift-ssb

Register a current clean physical 512 distribution with exact phase/loss parity, cache policy, and memory.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

CoM, DPC, and iDPC

rotation and Fourier iDPC from resident CoM maps

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident float32 CoM row and column maps

synchronized rotation and iDPC publication

3

dpc-parity

Repeat the complete CoM row/column, rotation, and iDPC suite with full-precision parity, synchronized timing, and memory telemetry.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared exact resident input

complete exact resident volume and seven products

1

swift-indexed-load

Repeat uncontended full-native publication with complete product hashes, p50/p95/max, driver allocation, RSS, pressure, and swap.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused

exact indexed load, detector sum, and seven native-resolution products

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused

exact indexed load, detector sum, and seven native-resolution products

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

Not supported

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

not applicable

unsupported load-plan contract

4

swift-indexed-load

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized display-ready outputs

3

display-fft-parity

Run the current Metal display-kernel distribution with exact revision, full parity, first-versus-warm timing, and peak memory.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident Metal maps

synchronized Fourier output

3

display-fft-parity

Run the current Metal FFT distribution with exact revision, full parity, first-versus-warm timing, and peak memory.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

cold original compressed-source full-native load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold original source; no prepared index

source selection through exact complete resident output and products

1

swift-indexed-load

Run arbitrary-source cold discovery, audit, index, full-native load, products, and complete memory/page-fault telemetry on physical MacBook Pro (M5, 24 GB).

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

I/O and load

prepared-index full-native exact load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE descriptors; exact private destination; pressure gate stopped repetition

exact indexed load and full-scan products; app and separately timed hashes excluded

1

swift-indexed-load

Eliminate or safely bound the observed swapout before attempting a repeated full-native distribution; preserve the exact full-volume and product contract.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

cold original compressed-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold original source; no prepared index

source selection through exact complete working output and products

1

swift-indexed-load

Run arbitrary-source cold discovery, audit, index creation, exact sum-bin-2 load, products, and complete memory/page-fault telemetry.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-2 load with destination reuse

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private destination reused

exact indexed detector sum-bin-2 load and native-resolution products; app and separately timed hashes excluded

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

cold original compressed-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold original source; no prepared index

source selection through exact complete working output and products

2

swift-indexed-load

Run arbitrary-source cold discovery, audit, index creation, exact sum-bin-4 load, products, and complete memory/page-fault telemetry.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-4 fused load with destination reuse

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

2

swift-indexed-load

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete 200-trial plus Nelder-Mead fit

2

swift-ssb

Run the current full fit with fitted-parameter parity, deterministic repeatability, cache policy, memory, and timing.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

○ Pending

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared BF evidence

complete synchronized object, phase, and loss

2

swift-ssb

Run the current native 512 workflow with exact phase/loss parity, cache policy, memory, and repeat distributions.

WebGPU

MacBook Air (M2, 8 GB)

○ Pending

CoM, DPC, and iDPC

DPC row/column and optimized-rotation iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident exact hardware-WebGPU state

awaited GPU readback for each published product

2

dpc-parity

Run the physical hardware-browser path with frozen full-precision parity, awaited GPU timing, pressure, swap, browser-tree memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admissible

blocked exact-residency contract

4

detector-product-parity

Implement and prove an exact bounded-streaming browser product path that never requires an 18 GiB resident input before attempting this physical row.

WebGPU

MacBook Air (M2, 8 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized display-ready publication

3

display-fft-parity

Run first-execution and warm hardware-browser display-kernel distributions with parity, presentation timing, pressure, swap, browser-tree memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized Fourier publication

3

display-fft-parity

Run first-execution and warm hardware-browser FFT distributions with parity, presentation timing, pressure, swap, browser-tree memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

full-native resident admission

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admitted on the 8 GB resident path

admission decision before allocation

1

webgpu-load

Keep this fail-closed until exact bounded residency can preserve the full logical tensor without an 18 GiB browser/device destination.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete 4.5 GiB working output

1

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

prepared-index exact detector-bin-2 hardware load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact complete 4.5 GiB working output

1

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete working output

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

prepared-index exact detector-bin-4 hardware load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact complete working output

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-8 load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete working output

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

prepared-index exact detector-bin-8 hardware load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact complete working output

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Air (M2, 8 GB)

○ Pending

Selective loading

64 by 64 prepared-shard rectangle

64 × 64

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes and frame-span manifest

exact selected resident output; harness setup excluded

2

webgpu-selective-load

Run the physical browser selector with full selected-tensor parity, source bytes, pressure, swap, browser/device memory, and timing; retain the intra-shard range-read gap explicitly.

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

CoM, DPC, and iDPC

DPC row/column and optimized-rotation iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident exact full-native WebGPU state

awaited GPU readback for each published product

3

dpc-parity

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact hardware-WebGPU input; prepared QH5 indexes and warm source pages declared by measurement

complete internal GPU product suite through awaited readbacks; parity hashing excluded

2

detector-product-parity

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized display-ready publication

3

display-fft-parity

Add a hardware-browser display-kernel entry point with parity, presentation, browser-tree memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized Fourier publication

3

display-fft-parity

Add a hardware-browser FFT entry point with parity, presentation, browser-tree memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

cold original compressed-source hardware load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete resident output

1

webgpu-load

Run seven independently controlled arbitrary-source cold browser loads with full parity and memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

1

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-2 hardware load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume; no float32 resident substitution

1

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

prepared-index hardware-WebGPU full-native load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes; warm source pages; fresh browser target per retained run

navigation through exact full-native resident output and diagnostic frame checksums

2

webgpu-load

Add total WebGPU device-allocation telemetry when the runtime exposes it; keep cold-source and application-E2E trials as separately labeled gates. Repeated harness wall is not yet at or below one second.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-4 hardware load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume; no float32 resident substitution

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

native uint16 exact residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

source and prepared state declared by the run

complete exact resident output

3

webgpu-load

Add WebGPU device-allocation telemetry and retain an exact complete-volume hash inside the timed boundary before promoting the integrated load gate.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

native uint32 exact residency

512 × 512

192 × 192

1

192 × 192

uint32

uint32

uint32

exact

source and prepared state declared by the run

complete exact resident output

3

webgpu-load

Run a native uint32 HDF5 source on physical hardware WebGPU and retain the complete resident hash, browser/device memory, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

native uint8 exact residency

512 × 512

192 × 192

1

192 × 192

uint8

uint8

uint8

exact

source and prepared state declared by the run

complete exact resident output

3

webgpu-load

Run an exact physical hardware-browser full-volume uint8 source through local HDF5 decode and retain output hash, allocation, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

audited lossless uint16 to uint8 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint8

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

3

registry-smoke

Replace the experimental low8 global flag with a typed source-identity-bound audit in the local-HDF5 public result, then retain exact full-volume parity.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-8 load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-8 hardware load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume; no float32 resident substitution

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

audited uint8 staging to exact uint16 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

Audited uint8 staging followed by exact uint16 resident reconstruction is not implemented in the public WebGPU local-HDF5 path.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

I/O and load

exact uint16 to uint32 widening

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint32

exact

source and prepared state declared by the run

complete exact widened resident output

4

registry-smoke

The WebGPU local-HDF5 public path does not widen uint16 source values into uint32 resident storage.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

I/O and load

explicit saturating uint16 to uint8 browse residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint8

browse-only

source and prepared state declared by the run

complete saturating browse resident output

4

webgpu-load

Run the fused full-plane clip8 path on physical hardware with values above 255 and retain output hash, browser/device memory, and browse-only provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

Screening

prepared screening workflow

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not applicable

unsupported contract test

5

registry-smoke

The backend parity contract currently marks the prepared screening workflow not implemented for WebGPU.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Selective loading

64 by 64 prepared-shard rectangle

64 × 64

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes and frame-span manifest

exact selected resident output; harness setup excluded

2

webgpu-selective-load

Add full selected-tensor parity and true intra-shard range reads; current evidence uses qualified frame probes and shard-level selection.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

Selective loading

arbitrary evidence-selective positions with order and duplicates

n/a

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source-page state declared

exact ordered selected resident output

2

webgpu-selective-load

Implement ordered/duplicate sparse selectors and true range reads, then run full physical parity, bytes, memory, and timing.

WebGPU

MacBook Pro (M5 Max, 128 GB)

× Refuted

Single-sideband ptychography

object, phase, and loss

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared detector-bin-2 evidence

complete physical browser object, phase, and loss

1

webgpu-ssb-preflight

Restore frozen phase parity before any performance promotion; retain the current mismatch as refuted.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

Single-sideband ptychography

200 trials plus Nelder-Mead

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

not applicable

unsupported contract test

5

registry-smoke

The frozen backend contract does not implement WebGPU calibration fitting.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

CoM, DPC, and iDPC

DPC row/column and optimized-rotation iDPC

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident exact hardware-WebGPU state

awaited GPU readback for each published product

2

dpc-parity

Run the physical hardware-browser path with frozen full-precision parity, awaited GPU timing, browser-tree memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact full-native hardware-WebGPU input

complete product suite publication

1

detector-product-parity

Run a guarded physical hardware-browser admission smoke, then retain the complete native product suite with full parity, timing, browser/device memory, pressure, and swap if admission remains safe.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

1

detector-product-parity

Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

Display kernels

transform, histogram, and color

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized display-ready publication

3

display-fft-parity

Run first-execution and warm hardware-browser display-kernel distributions with parity, presentation timing, browser-tree memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

FFT

Fourier transform

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident hardware-WebGPU maps

synchronized Fourier publication

3

display-fft-parity

Run first-execution and warm hardware-browser FFT distributions with parity, presentation timing, browser-tree memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

cold original compressed-source full-native hardware load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete resident output

1

webgpu-load

After bounded admission passes, run controlled arbitrary-source cold browser trials with full parity and complete memory/page-fault telemetry.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-2 load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

1

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

prepared-index exact full-native hardware load

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact full-native resident output

2

webgpu-load

Run one bounded admission smoke before a repeated exact production-path distribution with browser-tree and adapter memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-2 hardware load

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-4 load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

2

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-4 hardware load

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

cold original compressed-source exact detector-bin-8 load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

cold original source; no prepared index

file selection through exact complete uint16 working output

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run controlled arbitrary-source cold trials with full-volume parity, stage timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

I/O and load

prepared-index exact integer detector-bin-8 hardware load

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

prepared block indexes; source-page state declared by run

exact uint16 working volume

3

webgpu-load

Integrate and qualify exact-integer accumulation and residency, then run full-volume parity, repeated prepared timing, adapter/browser memory, and device-loss checks.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

Selective loading

64 by 64 prepared-shard rectangle

64 × 64

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes and frame-span manifest

exact selected resident output; harness setup excluded

2

webgpu-selective-load

Run the physical browser selector with full selected-tensor parity, source bytes, browser/device memory, and timing; retain the intra-shard range-read gap explicitly.

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact full-native CPU input

complete reference product suite

4

detector-product-profile

Run the independent full-native CPU reference product suite only when host CPU, storage, and memory ownership are uncontended; retain all hashes, timing, and peak RSS.

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

I/O and load

independent full-native exact load reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source-page state declared by run; CPU exact-reference validation

public CPU reference load; full-volume hash excluded

4

python-load-matrix

Run only when host CPU, storage, and memory ownership are uncontended; retain exact full-volume parity, p50/p95/max, RSS, cache state, and source bytes.

CPU reference

MacBook Air (M2, 8 GB)

! Blocked

Detector products

mean DP, total, BF, ABF, ADF, and DF reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admissible

blocked reference product suite

4

detector-product-profile

Implement and prove an exact bounded-streaming CPU reference product path that never requires an 18 GiB resident input before attempting this physical row.

CPU reference

MacBook Air (M2, 8 GB)

! Blocked

I/O and load

full-native exact CPU-reference admission

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

not admitted as an in-memory CPU result on the 8 GB computer

admission decision before allocation

4

python-load-matrix

Keep full resident admission fail-closed; qualify only an exact bounded mmap or streaming reference that preserves all 18 GiB logical bytes without memory-pressure termination.

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

exact full-native CPU-resident input

complete reference product suite

4

detector-product-profile

Add ABF, retain all product hashes, and repeat the exact physical CPU reference traversal for p50/p95/maximum and complete process-memory telemetry.

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

independent full-native exact load reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source pages unspecified; CPU exact-reference validation only

public CPU reference load; full-volume hash excluded

4

python-load-matrix

Keep the canonical CPU reference current after loader or metadata changes; a repeated CPU speed distribution is not a release gate.

CPU reference

MacBook Pro (M5, 24 GB)

○ Pending

Detector products

mean DP, total, BF, ABF, ADF, and DF reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident exact full-native CPU input

complete reference product suite

4

detector-product-profile

Run only after an uncontended memory preflight; retain one exact full-native CPU reference smoke, complete product hashes, pressure, swap, and peak RSS before repetition.

CPU reference

MacBook Pro (M5, 24 GB)

○ Pending

I/O and load

independent full-native exact load reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared indexes; source-page state declared by run; CPU exact-reference validation

public CPU reference load; full-volume hash excluded

4

python-load-matrix

Run one bounded exact CPU-reference smoke with mmap or streaming, full-volume parity, process physical footprint, pressure, and swap; repeat only if the 18 GiB logical output is safely admitted.

CPU reference

Portable CI runner

◐ Partial

CoM, DPC, and iDPC

CoM row/column, rotation, and iDPC reference

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident reference inputs

complete reference publication

4

dpc-parity

Maintain frozen full-precision references and signal-range checks.

CPU reference

Portable CI runner

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

resident reference input

complete reference product suite

4

detector-product-profile

Keep exact reference tests current; no CPU speed claim is required.

CPU reference

Portable CI runner

◐ Partial

Display kernels

transform, histogram, and color reference

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident reference maps

complete reference display outputs

4

display-fft-parity

Maintain the shared transform, histogram, and color goldens and frozen display conventions.

CPU reference

Portable CI runner

◐ Partial

FFT

Fourier transform reference

512 × 512

n/a

n/a

n/a

float32 maps

n/a

float32

n/a

resident reference maps

complete reference Fourier output

4

display-fft-parity

Maintain the frozen Fourier-transform reference and numerical conventions.

CPU reference

Portable CI runner

Not supported

I/O and load

audited uint8 staging to exact uint16 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

Audited uint8 staging with uint16 resident reconstruction is not implemented for CPU reference.

CPU reference

Portable CI runner

◐ Partial

I/O and load

native uint16 exact residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

Promote only after the portable public API fixture retains complete-volume uint16 hash, shape, order, mask, and metadata parity.

CPU reference

Portable CI runner

◐ Partial

I/O and load

native uint32 exact residency

512 × 512

192 × 192

1

192 × 192

uint32

uint32

uint32

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

Add a public-API native uint32 fixture that disables advisory auto narrowing and proves complete count, dtype, shape, and metadata parity.

CPU reference

Portable CI runner

◐ Partial

I/O and load

native uint8 exact residency

512 × 512

192 × 192

1

192 × 192

uint8

uint8

uint8

exact

source and prepared state declared by the run

complete exact resident output

4

registry-smoke

Add a repository fixture that proves public io.load preserves every native uint8 count, shape, order, and provenance field at detector bin 1.

CPU reference

Portable CI runner

Not supported

I/O and load

exact uint16 to uint32 widening

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint32

exact

source and prepared state declared by the run

complete exact widened resident output

4

registry-smoke

CPU reference io.load has no uint16-to-uint32 widening.

CPU reference

Portable CI runner

Not supported

I/O and load

audited lossless uint16 to uint8 residency

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint8

exact

source identity and complete value-range audit declared by the run

complete audit-bound exact resident output

4

registry-smoke

CPU reference io.load has no uint8 output dtype.

CPU reference

Portable CI runner

Not supported

I/O and load

explicit saturating uint16 to uint8 browse residency

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint8

browse-only

source and prepared state declared by the run

complete saturating browse resident output

4

registry-smoke

CPU reference io.load has no saturating uint8 browse output.

CPU reference

Portable CI runner

◐ Partial

I/O and load

independent exact load/bin reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

portable reference; performance not promoted

exact reference output

4

registry-smoke

Retain a standardized full-size CPU reference timing only when a practical physical memory envelope is available; parity remains the primary role.

CPU reference

Portable CI runner

◐ Partial

Screening

complete screening reference

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

reference source

complete reference screening result

4

screening-profile

Maintain the frozen reference arrays; physical timing is not a CPU release gate.

CPU reference

Portable CI runner

◐ Partial

Single-sideband ptychography

object, phase, and loss reference

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared reference BF evidence

complete reference object, phase, and loss

4

registry-smoke

Maintain the frozen full-precision reference fixture and objective.

CPU reference

Portable CI runner

◐ Partial

Single-sideband ptychography

200 trials plus Nelder-Mead reference

512 × 512

n/a

n/a

n/a

full active BF evidence

n/a

complex64

n/a

prepared reference BF evidence

complete 200-trial plus Nelder-Mead fit

5

registry-smoke

Maintain the frozen deterministic reference objective and fitted parameters.

Retained atomic measurements#

These rows are imported from immutable evidence or explicitly registered follow-up evidence. A partial row is retained but does not satisfy a complete timing-and-parity gate.

Platform

Computer

State

Module

Operation

Selected scan

Source detector

Detector bin

Output detector

Source dtype

Staging dtype

Resident dtype

Scientific gate

Cache/process state

Wall boundary

Samples

p50

p95

Maximum

Logical resident

Driver allocated after load

Driver allocated after release

Accelerator peak

Total-device peak

Process/tree peak

Process physical-footprint peak

Swap delta

Parity

Device tested

Date tested

Revision

Fixture ID

Master SHA-256

Source identity SHA-256

Measurement ID

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-streamed-summaries

n/a

warm operating-system source pages; unspecified; no prepared source index

quantem.gpu.screening.prepare package wall through exact complete public screening arrays

3

1.208146 s

1.241817 s

1.245558 s

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

5ee2016e21528bcc05bf26801c3ff6b244359e21

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-contiguous-read-plan

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-streamed-summaries

n/a

warm operating-system source pages; unspecified; no prepared source index

quantem.gpu.screening.prepare package wall through exact complete public screening arrays

3

1.663968 s

1.718105 s

1.724120 s

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

5ee2016e21528bcc05bf26801c3ff6b244359e21

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-exact-mask-guard

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-streamed-summaries

n/a

warm operating-system source pages; unspecified; no prepared source index; unrelated CPU contention retained

quantem.gpu.screening.prepare package wall through exact complete public screening arrays

3

1.546580 s

1.552670 s

1.553347 s

n/a

n/a

n/a

n/a

6.084 GiB

2.130 GiB

n/a

n/a

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

5ee2016e21528bcc05bf26801c3ff6b244359e21

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-final-source-distribution

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-streamed-summaries

n/a

warm operating-system source pages; unspecified; empty screening-result cache

quantem.gpu.screening.prepare package wall through exact complete six-array screening result in balanced A-B-B-A-B-A-A-B comparison

6

1.356516 s

1.642386 s

1.730328 s

n/a

n/a

n/a

2.084 GiB

6.084 GiB

2.130 GiB

n/a

0 B

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

023a6c497b106b216c87205d3fbec63377d77177

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-pinned-slot-baseline

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-streamed-summaries

n/a

warm operating-system source pages; unspecified; empty screening-result cache

quantem.gpu.screening.prepare package wall through exact complete six-array screening result in balanced A-B-B-A-B-A-A-B comparison

6

1.204713 s

1.325760 s

1.329731 s

n/a

n/a

n/a

2.084 GiB

6.084 GiB

1.958 GiB

n/a

0 B

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

023a6c497b106b216c87205d3fbec63377d77177

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-pinned-slot-candidate

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓ Measured

Screening

screening-cache-reopen

512 × 512

192 × 192

1

192 × 192

uint16

n/a

not-applicable-cached-results

n/a

prepared immutable screening result cache; fresh-process reopen; source build excluded

fresh-process exact-spelling prepared-result reopen through strong source validation, lazy raw-I/O import boundary, six-array materialization, and immutable ScreeningResult publication

3

5.599 ms

5.771 ms

5.790 ms

n/a

n/a

n/a

n/a

n/a

0.874 GiB

n/a

0 B

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0

2026-08-22

8258ac2c62f9cc3c2739027352411df3bb3eae19

real-512-native-detector

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

linux-dual-blackwell-96gb-cuda-result-cache-final-reopen

Python MPS

MacBook Pro (M5 Max, 128 GB)

× Refuted

CoM, DPC, and iDPC

CoM row/column, rotation, aligned fields, and iDPC parity smoke

512 × 512

192 × 192

1

192 × 192

uint16

uint16

float32 products

full-precision

exact resident detector-bin-1 MPS input; source pages uncontrolled or warm; no prepared product result; five pageouts; zero swapout growth

single first-execution diagnostic only; synchronized public CoM through rotation and iDPC pipeline was 0.159022 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

n/a

27.510 GiB

n/a

20.436 GiB

28.001 GiB

0 B

Failed

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b8df61b55920cb098e2848301f6d30c45870511d

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-python-mps-bin1-dpc-idpc-refuted-b8df61b

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Detector products

mean DP, total, BF, ABF, ADF, and DF exact parity smoke

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16 source; uint64 and float32 products

full-precision

exact resident detector-bin-1 MPS input; source pages uncontrolled or warm; no prepared product result; five pageouts; zero swapout growth

single first-execution diagnostic only; synchronized detector-product publication was 0.499909 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

n/a

27.510 GiB

n/a

20.436 GiB

28.001 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b8df61b55920cb098e2848301f6d30c45870511d

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-python-mps-bin1-detector-products-smoke-b8df61b

Python MPS

MacBook Pro (M5 Max, 128 GB)

× Refuted

FFT

first-execution Torch-MPS FFT parity smoke

512 × 512

192 × 192

1

192 × 192

uint16

uint16

float32

full-precision

resident full-precision iDPC map; first MPS FFT execution; five pageouts; zero swapout growth

single first-execution diagnostic only; upload, synchronized FFT, and readback publication was 0.291350 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

n/a

27.510 GiB

n/a

20.436 GiB

28.001 GiB

0 B

Failed

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b8df61b55920cb098e2848301f6d30c45870511d

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-python-mps-bin1-fft-refuted-b8df61b

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

post-warmup full-native exact load with fresh returned destination

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

one same-process warmup; operating-system source pages uncontrolled; no eviction performed; fresh returned destination released after each trial

public io.load return after MPS synchronization; full-volume hash and release excluded

7

0.406624 s

0.428164 s

0.428164 s

18.000 GiB

18.442 GiB

0.442 GiB

18.442 GiB

n/a

18.692 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-python-mps-bin1-canonical-clean-r7

Python MPS

MacBook Pro (M5 Max, 128 GB)

↺ Superseded

I/O and load

post-warmup exact package load with fresh returned destination

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

one same-process warmup; operating-system source pages uncontrolled; no eviction performed; fresh returned destination released after each trial

public io.load return after backend synchronization

7

0.414824 s

0.457261 s

0.457261 s

18.000 GiB

18.442 GiB

0.442 GiB

n/a

n/a

0.691 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

0bc9378b15285db4de69baa399cfc573936a2e41

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-python-mps-bin1-final-head-fresh-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; fresh destination

exact package load into a fresh destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.425533 s

0.436353 s

0.437419 s

18.000 GiB

n/a

n/a

18.442 GiB

n/a

0.687 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

e8293a1f2b9eef6ddccd2ed06ab6820008480e82

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin1-fresh-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; explicit resident destination recycled

exact package load into caller-selected resident destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.259189 s

0.263118 s

0.263375 s

18.000 GiB

n/a

n/a

18.442 GiB

n/a

0.688 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

e8293a1f2b9eef6ddccd2ed06ab6820008480e82

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin1-recycled-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; fresh destination

exact package load into a fresh caller-selected resident destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.462541 s

0.479014 s

0.483058 s

4.500 GiB

n/a

n/a

5.688 GiB

n/a

0.571 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b7f8ef3ff2a2d8458944e1e55a3296a39c854357

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin2-fresh-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; explicit resident destination recycled

exact package load into an explicitly recycled caller-selected resident destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.359606 s

0.361384 s

0.361995 s

4.500 GiB

n/a

n/a

5.688 GiB

n/a

0.571 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b7f8ef3ff2a2d8458944e1e55a3296a39c854357

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin2-recycled-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

loading

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

warm operating-system source pages; uncontrolled

fresh-process package load only; not import, cold HDF5, prepared reopen, first app product, or app E2E

14

0.498708 s

0.503702 s

0.505150 s

4.500 GiB

n/a

n/a

5.688 GiB

n/a

0.575 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

341a0365175249773acec600137f2b00aa4482db

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin2-specialized

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; fresh destination

exact package load into a fresh caller-selected resident destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.384264 s

0.385355 s

0.385638 s

1.125 GiB

n/a

n/a

2.313 GiB

n/a

0.571 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b7f8ef3ff2a2d8458944e1e55a3296a39c854357

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin4-fresh-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

loading

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

warm or uncontrolled operating-system source pages after one same-lifecycle warmup; explicit resident destination recycled

exact package load into an explicitly recycled caller-selected resident destination in balanced A-B-B-A-B-A-A-B lifecycle comparison

8

0.352990 s

0.355048 s

0.355062 s

1.125 GiB

n/a

n/a

2.313 GiB

n/a

0.572 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b7f8ef3ff2a2d8458944e1e55a3296a39c854357

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin4-recycled-destination

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

loading

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared plans and warm or uncontrolled operating-system source pages

fresh-process exact package load after one excluded warm-up

6

0.383772 s

0.390006 s

0.391839 s

1.125 GiB

n/a

n/a

2.313 GiB

n/a

n/a

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

0a0d60eef298190597778dbf11bb5b79f97a3615

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin4-specialized

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

loading

512 × 512

192 × 192

8

24 × 24

uint16

n/a

uint16

n/a

warm operating-system source pages; uncontrolled

immutable-source fresh-process package A-B-B-A; contended stage run excluded

6

0.356969 s

0.359302 s

0.359820 s

0.281 GiB

n/a

n/a

1.470 GiB

n/a

0.572 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

38bf65ce33574513745b255fd15e8b3f1d52b6a9

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-python-mps-bin8-specialized

Python MPS

MacBook Pro (M5, 24 GB)

◐ Partial

CoM, DPC, and iDPC

CoM row/column, rotation, aligned fields, and iDPC parity smoke

512 × 512

192 × 192

2

96 × 96

uint16

n/a

float32 products

n/a

exact resident detector-bin-2 MPS input; source pages uncontrolled or warm; no prepared product result; Screen Sharing active; 73 pageouts; zero swap growth

single contaminated diagnostic only; public CoM through rotation and iDPC pipeline was 0.080079 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

4.500 GiB

n/a

n/a

5.736 GiB

n/a

9.657 GiB

10.508 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-python-mps-bin2-dpc-idpc-smoke-68dbe3a

Python MPS

MacBook Pro (M5, 24 GB)

× Refuted

FFT

first-execution Torch-MPS FFT parity smoke

512 × 512

192 × 192

2

96 × 96

uint16

n/a

float32

n/a

resident full-precision iDPC map; first MPS FFT execution; Screen Sharing active; 73 pageouts; zero swap growth

single contaminated diagnostic only; upload, synchronized FFT, and readback publication was 0.212618 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

4.500 GiB

n/a

n/a

5.736 GiB

n/a

9.657 GiB

10.508 GiB

0 B

Failed

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-python-mps-bin2-display-fft-refuted-68dbe3a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-2 load with destination reuse

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.225984 s

0.231192 s

0.231192 s

4.500 GiB

n/a

n/a

5.087 GiB

n/a

0.657 GiB

5.336 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-swift-metal-bin2-canonical-r7

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-4 load with destination reuse

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

6

0.200036 s

0.204694 s

0.204694 s

1.125 GiB

n/a

n/a

1.712 GiB

n/a

1.039 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-swift-metal-bin4-canonical-r7

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

↺ Superseded

I/O and load

loading-and-products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

fresh QH5 index root and destination; F_NOCACHE on source hashing and indexed descriptors; already-audited immutable source

catalog plus pipeline compilation, plan, complete exact private-resident volume, and seven products

7

0.524590 s

1.572781 s

1.572781 s

18.000 GiB

n/a

n/a

18.571 GiB

n/a

0.991 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GiB unified memory

2026-08-22

23d25619cfe22d5e89761fda2d2796a7c82ba090

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-swift-metal-controlled-fullnative

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

controlled uncached exact resident load and products

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

fresh QH5 index root and destination; macOS F_NOCACHE on source hashing and indexed descriptors; immutable source already audited

catalog, pipeline compilation, plan, complete exact private-resident volume, seven products, metadata, and provenance

7

0.577793 s

0.900979 s

0.900979 s

18.000 GiB

n/a

n/a

18.571 GiB

n/a

0.874 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

c0ea44465e6346a8436a0b74f491a04af0b5dc32

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-swift-metal-controlled-fullnative-current

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

I/O and load

loading-and-products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable QH5 indexes; source pages unspecified

complete exact resident volume, products, metadata, calibration, and provenance; full private readback/hash excluded

6

0.313870 s

0.318865 s

0.318865 s

18.000 GiB

n/a

n/a

18.571 GiB

n/a

0.639 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

531e5001f5a0a886c058f772ef42d770f000890b

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-max-128gb-swift-metal-private-bin1-load

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

CoM, DPC, and iDPC

dpc-idpc

512 × 512

192 × 192

1

192 × 192

uint16

n/a

float32

n/a

already-resident 512x512 float32 CoM row and column maps; source load excluded

explicit CoM rotation plus Fourier iDPC through synchronized publication; source load and CoM reduction excluded

15

1.979 ms

3.640 ms

3.640 ms

n/a

n/a

n/a

0.009 GiB

n/a

0.034 GiB

0.064 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

0b305f69fc1039933463986a0f74609e22d4dd35

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-24gb-swift-metal-resident-dpc-idpc

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

original HDF5 reread to complete exact packed resident

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

exact

original compressed source reread every visit; existing QH5 index and source-bound packing layout; exact DPC sums reused after first load; OS pages uncontrolled

indexed source open through synchronous complete packed-resident return; catalog and independent full-count audit separate

7

1.496660 s

2.086119 s

2.086119 s

18.000 GiB

n/a

n/a

n/a

n/a

n/a

2.531 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-09-06

e305f9216ed359397e27f69882931ffe16de8d99

original-512x512x192x192-u16-fixture-d

b364e25625ca9749ad5251eed4ebe28a771d727073f361ae35eb7cd42443117e

532d00ffbb084bf32c96f5a53cde99e9085e91f0490e925c530d92bfdba4b350

macbook-pro-m5-24gb-m5-original-hdf5-packed-current

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

I/O and load

prepared-index full-native exact private-resident smoke

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE descriptors; fresh exact private-resident destination; repetition stopped by pressure gate

exact indexed full-native load and seven native-resolution products; source discovery, audit creation, application presentation, and separately timed hashes excluded

1

3.424108 s

3.424108 s

3.424108 s

18.000 GiB

n/a

n/a

18.587 GiB

n/a

0.611 GiB

18.673 GiB

~0.706 GiB

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin1-canonical-smoke

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

I/O and load

loading-and-products

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified

package exact indexed load plus exact full-scan products; app UI, index creation, and full-volume hash excluded

1

n/a

n/a

n/a

18.000 GiB

n/a

n/a

18.571 GiB

n/a

0.624 GiB

18.658 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

0b305f69fc1039933463986a0f74609e22d4dd35

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-24gb-swift-metal-bin1-single-admission

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-2 load with destination reuse

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.678703 s

0.697179 s

0.697179 s

4.500 GiB

n/a

n/a

5.087 GiB

n/a

0.662 GiB

5.271 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin2-canonical-r7

Native Swift/Metal

MacBook Pro (M5, 24 GB)

↺ Superseded

I/O and load

prepared-index exact detector-bin-2 canonical smoke

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; fresh exact private-resident destination

catalog, pipeline compilation, plan, exact indexed detector sum-bin-2 load, and seven native-resolution products; separately timed boundary hashes excluded

1

0.763019 s

0.763019 s

0.763019 s

4.500 GiB

n/a

n/a

5.087 GiB

n/a

0.645 GiB

5.200 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin2-canonical-smoke

Native Swift/Metal

MacBook Pro (M5, 24 GB)

↺ Superseded

I/O and load

prepared-index exact detector-bin-2 load with destination reuse

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, first allocation, pipeline compilation, and application excluded

6

0.671309 s

0.694479 s

0.694479 s

4.500 GiB

n/a

n/a

5.071 GiB

n/a

0.646 GiB

n/a

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

75be74e899186dfb86f08f0c591de8c65f0c5e53

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin2-current-head-reuse

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

loading-and-products

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified; caller destination reused

package exact indexed load plus exact full-scan products; hash and app excluded

6

0.683492 s

0.690893 s

0.690893 s

4.500 GiB

n/a

n/a

5.071 GiB

n/a

0.647 GiB

5.182 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

0b305f69fc1039933463986a0f74609e22d4dd35

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-24gb-swift-metal-bin2-reuse

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

prepared-index exact detector-bin-4 load with destination reuse

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.640942 s

0.651931 s

0.651931 s

1.125 GiB

n/a

n/a

1.712 GiB

n/a

1.018 GiB

2.319 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin4-canonical-r7

Native Swift/Metal

MacBook Pro (M5, 24 GB)

↺ Superseded

I/O and load

prepared-index exact detector-bin-4 load with destination reuse

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, first allocation, pipeline compilation, and application excluded

6

0.630903 s

0.638497 s

0.638497 s

1.125 GiB

n/a

n/a

1.696 GiB

n/a

0.650 GiB

n/a

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

75be74e899186dfb86f08f0c591de8c65f0c5e53

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-24gb-swift-metal-bin4-current-head-reuse

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

I/O and load

loading-and-products

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified; caller destination reused

package load plus seven exact products; UI and full-volume hash excluded

12

0.631540 s

0.645521 s

0.645521 s

1.125 GiB

n/a

n/a

1.696 GiB

n/a

0.654 GiB

n/a

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

34782081d43f0dcefe6fd13c9d80d1e39dfc505b

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

n/a

macbook-pro-m5-24gb-swift-metal-bin4-fused-reuse

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

CoM, DPC, and iDPC

dpc

512 × 512

192 × 192

1

192 × 192

uint16

n/a

float32

n/a

resident exact full-native WebGPU state; source load excluded

DPC column through awaited GPU readback; synchronized wall time, not GPU-only timestamp interval

7

0.700 ms

0.800 ms

0.800 ms

n/a

n/a

n/a

n/a

n/a

6.366 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

64b6eec675e566cc5b4f31343dc660788f10ee48

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

macbook-pro-m5-max-128gb-webgpu-resident-dpc-column

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

CoM, DPC, and iDPC

dpc

512 × 512

192 × 192

1

192 × 192

uint16

n/a

float32

n/a

resident exact full-native WebGPU state; source load excluded

DPC row through awaited GPU readback; synchronized wall time, not GPU-only timestamp interval

7

0.700 ms

0.900 ms

0.900 ms

n/a

n/a

n/a

n/a

n/a

6.366 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

64b6eec675e566cc5b4f31343dc660788f10ee48

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

macbook-pro-m5-max-128gb-webgpu-resident-dpc-row

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

CoM, DPC, and iDPC

idpc

512 × 512

192 × 192

1

192 × 192

uint16

n/a

float32

n/a

resident exact full-native WebGPU state; source load excluded

optimized-rotation iDPC through awaited GPU readback; synchronized wall time, not GPU-only timestamp interval

7

1.400 ms

1.500 ms

1.500 ms

n/a

n/a

n/a

n/a

n/a

6.366 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

64b6eec675e566cc5b4f31343dc660788f10ee48

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

macbook-pro-m5-max-128gb-webgpu-resident-idpc-optimized

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

Detector products

exact resident full detector-product plus CoM/DPC/iDPC suite

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared QH5 indexes; warm source pages; exact full-native uint16 resident input; fresh browser target per retained run

complete internal GPU product suite from exact resident input through awaited readbacks; parity hashing, sorting, and encoding excluded

7

0.482000 s

0.483600 s

0.483600 s

18.000 GiB

n/a

n/a

n/a

n/a

6.492 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-fullnative-bin1-products-r7-d8e6f56

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

prepared-index full-native exact hardware-WebGPU distribution

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared immutable block indexes; explicitly warm source pages; fresh browser target per retained run; 699 pageouts; zero swap growth; no dropped runs

navigation through scientifically usable exact resident output and diagnostic frame checksums; exhaustive full-volume hash and application E2E excluded

7

1.358000 s

1.594000 s

1.594000 s

18.000 GiB

n/a

n/a

n/a

n/a

6.500 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-fullnative-exact-r7-d8e6f56

WebGPU

MacBook Pro (M5 Max, 128 GB)

↺ Superseded

I/O and load

prepared-index full-native exact hardware-WebGPU smoke

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable block indexes; source pages warm from the immediately preceding independent CPU reference; fresh hardware browser target

hardware harness load wall through exact resident output; 63.153-second full-volume verification hash and application E2E excluded

1

1.094000 s

1.094000 s

1.094000 s

18.000 GiB

n/a

n/a

n/a

n/a

≥6.604 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

34f0029c9fe3b92c2709283547301a91899e4d66

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-fullnative-exact-smoke-34f0029

WebGPU

MacBook Pro (M5 Max, 128 GB)

↺ Superseded

I/O and load

loading

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared block indexes; source pages unspecified and warm after the accepted parity run

one full-native hardware-WebGPU load preceding the accepted synchronized resident-product repetitions

1

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

6.366 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

64b6eec675e566cc5b4f31343dc660788f10ee48

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

n/a

macbook-pro-m5-max-128gb-webgpu-fullnative-load-single

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Selective loading

selective-loading

256 × 256

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable block indexes and frame-span manifest; operating-system source-page state uncontrolled/unspecified; no eviction performed; fresh browser target per repetition

WebGPU loader wall through exact selected resident output; frame-manifest read/encoding, DevTools injection, checksum harness, products, and application E2E excluded

5

0.381000 s

0.392400 s

0.394000 s

4.500 GiB

n/a

n/a

n/a

n/a

2.797 GiB

n/a

0 B

Qualified probes

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal WebGPU adapter; software=false

2026-08-22

23d25619cfe22d5e89761fda2d2796a7c82ba090

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-selective-256x256

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Selective loading

selective-loading

384 × 384

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable block indexes and frame-span manifest; operating-system source-page state uncontrolled/unspecified; no eviction performed; fresh browser target per repetition

WebGPU loader wall through exact selected resident output; frame-manifest read/encoding, DevTools injection, checksum harness, products, and application E2E excluded

5

0.574000 s

0.582000 s

0.584000 s

10.125 GiB

n/a

n/a

n/a

n/a

3.629 GiB

n/a

0 B

Qualified probes

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal WebGPU adapter; software=false

2026-08-22

23d25619cfe22d5e89761fda2d2796a7c82ba090

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-selective-384x384

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Selective loading

selective-loading

64 × 64

192 × 192

1

192 × 192

uint16

n/a

uint16

n/a

prepared immutable block indexes and frame-span manifest; operating-system source-page state uncontrolled/unspecified; no eviction performed; fresh browser target per repetition

WebGPU loader wall through exact selected resident output; frame-manifest read/encoding, DevTools injection, checksum harness, products, and application E2E excluded

5

0.147000 s

0.154400 s

0.156000 s

0.281 GiB

n/a

n/a

n/a

n/a

1.606 GiB

n/a

0 B

Qualified probes

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal WebGPU adapter; software=false

2026-08-22

23d25619cfe22d5e89761fda2d2796a7c82ba090

full-native-webgpu-512x512x192x192-u16

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-webgpu-selective-64x64

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

Detector products

independent mean DP, total, BF, ADF, DF, CoM, DPC, and iDPC reference traversal

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

exact full-native CPU-resident input after a warm operating-system page-cache source load

single independent detector-product traversal was 31.078040 seconds; source load excluded; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

n/a

n/a

n/a

36.450 GiB

n/a

n/a

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-19

334b7b5135fe29787540370a00f280fa138430a2

real-512x512x192x192-u16-bslz4-27shard

4802ec16ba241fef439e9dcb1c28e94f9cf9d95f773df9c5c8c3b5f7ed8192c4

1be810b96fdff8e384ad4cb6ebd49adff9b4ab0a6503cd5fed9106e09f5aa286

macbook-pro-m5-max-128gb-cpu-reference-bin1-products-334b7b5

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

I/O and load

independent full-native exact load reference

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared indexes; source pages unspecified; CPU exact-reference validation only

public CPU reference load; full-volume hash excluded

1

32.788039 s

32.788039 s

32.788039 s

18.000 GiB

n/a

n/a

n/a

n/a

19.141 GiB

n/a

0 B

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

real-512x512x192x192-u16-bslz4-27shard-master-fixture-c

c9c0d968fae70b8911ae925d676b90007886970fa99fe296a47cfe07844bbfe9

9f0ddb932c631b63cb573c38d747fa41941ee585c5389d33bdafb4add962b768

macbook-pro-m5-max-128gb-cpu-reference-bin1-canonical-smoke

Reproducible runbooks#

Runbook

Owner

Tier

Evidence level

Command

Required artifacts

cuda-ssb-preflight

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

PR smoke

parity preflight

PYTHONPATH=src python -m pytest -q tests/hardware/cuda/test_ssb_cuda_128.py tests/contracts/test_ssb_backend_contract.py tests/contracts/test_ssb_persistence.py

pytest log; separate physical reconstruction/calibration timing artifact

detector-product-parity

Portable CI runner plus specified backend computer

PR smoke

parity preflight

PYTHONPATH=src python -m pytest -q tests/parity/test_integer_detector_products.py tests/parity/test_products_parity.py tests/hardware/cuda/test_cuda_virtual_image.py

pytest log; backend-specific physical timing from a separate runbook

detector-product-profile

Physical CPU, CUDA, or Apple MPS computer named by the gate

scheduled physical profile

physical performance

PYTHONPATH=src python scripts/benchmark_detector_products.py "$QGPU_MASTER" --backend "$QGPU_BACKEND" --computer "$QGPU_PUBLIC_COMPUTER" --fixture-id "$QGPU_FIXTURE_ID" --source-sha256 "$QGPU_MASTER_SHA256" --expected-volume-sha256 "$QGPU_LOGICAL_PIXEL_SHA256" --expected-storage-shape "$QGPU_STORAGE_SHAPE" --reference-npz "$QGPU_DETECTOR_REFERENCE_NPZ" --reference-sha256 "$QGPU_REFERENCE_NPZ_SHA256" --scan-shape "$QGPU_SCAN_SHAPE" --source-detector-shape "$QGPU_SOURCE_DETECTOR_SHAPE" --working-detector-shape "$QGPU_WORKING_DETECTOR_SHAPE" --source-dtype "$QGPU_SOURCE_DTYPE" --working-dtype "$QGPU_WORKING_DTYPE" --source-value-maximum "$QGPU_SOURCE_MAX" --source-value-maximum-basis "$QGPU_RANGE_AUDIT_SHA256" --det-bin {detector_bin} --center-row "$QGPU_CENTER_ROW" --center-column "$QGPU_CENTER_COLUMN" --bf-radius-pixels "$QGPU_BF_RADIUS" --pixel-mask "$QGPU_PIXEL_MASK_POLICY" --cache-state "$QGPU_CACHE_STATE" --warmup 1 --reps 7 --memory-sample-ms 10 --json-out "$QGPU_OUTPUT_JSON"

run-level synchronized timing and exact product hashes; source, full-volume, independent-reference, range-audit, mask, dtype, shape, and detector-bin identities; accelerator/driver/process peak memory plus system pressure, compressor, pageout, and swap telemetry; experiment manifest and RUNS.md row

display-fft-parity

Portable CI runner plus specified backend computer

PR smoke

parity preflight

PYTHONPATH=src python -m pytest -q tests/parity/test_display_parity.py tests/parity/test_display_fft_parity.py tests/parity/test_display_shared_goldens.py

pytest log; physical first-execution and warm timing artifact

dpc-parity

Portable CI runner plus specified backend computer

PR smoke

parity preflight

PYTHONPATH=src python -m pytest -q tests/parity/test_dpc_rotation_agreement.py tests/parity/test_products_parity.py tests/parity/test_display_fft_parity.py

pytest log; physical backend synchronized timing and parity artifact

mps-ssb

MacBook Pro (M5 Max, 128 GB)

scheduled physical profile

physical performance

PYTHONPATH=src python scripts/benchmark_ssb_mps_scaling.py "$QGPU_SOURCE" --calibration "$QGPU_CALIBRATION" --sizes 512 --scan-sampling-a "$QGPU_SCAN_SAMPLING_A" --trials 200 --pair-iters 12 --parity-pairs 20 --reference-engine "$QGPU_REFERENCE_ENGINE" --jsonl-out "$QGPU_OUTPUT_JSONL"

run-level reconstruction and fit JSONL; phase/loss parity and fitted parameters; allocator/driver/process/swap memory; experiment manifest and RUNS.md row

python-load-matrix

specified physical backend computer

scheduled physical profile

physical performance

PYTHONPATH=src python scripts/benchmark_hdf5_load.py "$QGPU_MASTER" --backend "$QGPU_BACKEND" --dtype "$QGPU_EXPECTED_OUTPUT_DTYPE" --det-bin {detector_bin} --reps 7 --warmup {warmup_count} --cache-state "$QGPU_CACHE_STATE" --fixture-id "$QGPU_FIXTURE_ID" --source-sha256 "$QGPU_SOURCE_SHA256" --expected-output-sha256 "$QGPU_EXPECTED_OUTPUT_SHA256" --expected-output-dtype "$QGPU_EXPECTED_OUTPUT_DTYPE" --expected-output-shape "$QGPU_EXPECTED_OUTPUT_SHAPE" --expected-scan-shape "$QGPU_EXPECTED_SCAN_SHAPE" --expected-source-detector-shape "$QGPU_EXPECTED_SOURCE_DETECTOR_SHAPE" --expected-working-detector-shape "$QGPU_EXPECTED_WORKING_DETECTOR_SHAPE" --expected-source-dtype "$QGPU_EXPECTED_SOURCE_DTYPE" --require-full-output-parity --memory-sample-ms 10 --json-out "$QGPU_OUTPUT_JSON"

run-level JSON with p50, p95, and maximum; exact output parity hashes; logical resident, sampled accelerator/process lower-bound peaks, lifetime process RSS, and swap telemetry; external process-tree, compressed-memory, and pageout telemetry when required by the physical gate; experiment manifest and RUNS.md row

registry-smoke

Portable CI runner

PR smoke

parity preflight

PYTHONPATH=src python -m pytest -q tests/infrastructure/test_profile_registry.py tests/infrastructure/test_benchmark_registry.py

pytest terminal log; registry validator output

screening-profile

Linux CUDA workstation (dual 96 GB Blackwell GPUs) or MacBook Pro (M5 Max, 128 GB)

scheduled physical profile

physical performance

PYTHONPATH=src python scripts/benchmark_screening.py "$QGPU_MASTER" --backend "$QGPU_BACKEND" --cache-dir "$QGPU_SCREENING_CACHE_DIR" --cache-state "$QGPU_CACHE_STATE" --fixture-id "$QGPU_FIXTURE_ID" --source-sha256 "$QGPU_SOURCE_SHA256" --reps 7 --warmup 1 --reopen-reps 7 --reference-hashes-json "$QGPU_REFERENCE_HASHES_JSON" --json-out "$QGPU_OUTPUT_JSON"

run-level timing and cache lifecycle JSON; mean DP, total, BF, DF, CoM, rotation, and iDPC parity; GPU allocation/reserve/total-card and host RSS/swap telemetry; experiment manifest and RUNS.md row

swift-indexed-load

MacBook Pro (M5 Max, 128 GB), MacBook Air (M2, 8 GB), or MacBook Pro (M5, 24 GB)

scheduled physical profile

physical performance

swift run -c release metal-4dstem-indexed-load-benchmark --input "$QGPU_MASTER" --cache-dir "$QGPU_INDEX_DIR" --output-dir "$QGPU_OUTPUT_DIR" --revision "$QGPU_REVISION" --all-bands --exact-working-audit "$QGPU_AUDIT_JSON" --detector-bin {detector_bin} --maximum-shard-bytes "$QGPU_MAX_SHARD_BYTES" --destination-storage private --iterations 7 --buffered-read-ahead-shards 2 --uncached-source-reads --boundary-working-volume-hashes --expected-logical-pixel-sha256 "$QGPU_EXPECTED_LOGICAL_PIXEL_SHA256"

benchmark JSON, physical-storage hashes, and canonical logical-pixel hashes; source and working geometry/dtype/provenance; Metal allocation, process RSS/footprint, pressure, and swap telemetry; experiment manifest and RUNS.md row

swift-original-packed-load

Physical Apple Silicon Mac

scheduled physical profile

physical performance

swift run -c release metal-original-hdf5-benchmark "$QGPU_INPUT" "$QGPU_INDEX_DIR" --repeats 7 --plan-directory "$QGPU_PLAN_DIR" --reuse-products --oracle "$QGPU_ORACLE_JSON"

Raw benchmark JSON lines including failures and release records; Independent full-count oracle and source/working dtype and geometry; Separate process/device memory, pressure and swap telemetry

swift-ssb

MacBook Pro (M5 Max, 128 GB) or MacBook Pro (M5, 24 GB)

scheduled physical profile

physical performance

swift run -c release metal-ssb-benchmark "$QGPU_METADATA_JSON" "$QGPU_BF_U8" "$QGPU_REFERENCE_PHASE_F32" 7 full 200 3

benchmark JSON; phase/loss/fitted-parameter parity; Metal and process memory telemetry; experiment manifest and RUNS.md row

webgpu-load

MacBook Pro (M5 Max, 128 GB), MacBook Pro (M5, 24 GB), or MacBook Air (M2, 8 GB)

scheduled physical profile

physical performance

python scripts/benchmark_webgpu_h5_browser.py --cdp "$QGPU_CDP" --html-url "$QGPU_EXPORT_URL" --fixture-dir "$QGPU_FIXTURE_DIR" --master-name "$QGPU_MASTER_NAME" --data-template "$QGPU_DATA_TEMPLATE" --block-index-template "$QGPU_BLOCK_INDEX_TEMPLATE" --data-files "$QGPU_DATA_FILES" --frames "$QGPU_FRAMES" --det-bin {detector_bin} --reps 7 --require-local-profile --require-integer-resident --expected-resident-dtype {resident_dtype} --expected-resident-shape "$QGPU_EXPECTED_RESIDENT_SHAPE" --checksum-json "$QGPU_CHECKSUM_JSON" --require-checksum-parity --expected-full-output-sha256 "$QGPU_EXPECTED_FULL_OUTPUT_SHA256" --require-full-output-parity --json-out "$QGPU_OUTPUT_JSON"

browser profile JSON with run-level timings; independent full-output SHA-256 plus diagnostic frame checksums; browser-tree RSS, adapter allocation when available, and swap telemetry; experiment manifest and RUNS.md row

webgpu-selective-load

MacBook Pro (M5 Max, 128 GB) or MacBook Air (M2, 8 GB)

scheduled physical profile

physical performance

python scripts/benchmark_webgpu_h5_browser.py --html-url "$QGPU_EXPORT_URL" --fixture-dir "$QGPU_FIXTURE_DIR" --source-scan-shape 512,512 --scan-region "$QGPU_SCAN_REGION" --det-bin {detector_bin} --reps 7 --frame-index-json "$QGPU_FRAME_INDEX_JSON" --require-local-profile --checksum-json "$QGPU_CHECKSUM_JSON" --json-out "$QGPU_OUTPUT_JSON"

browser timing and source-byte JSON; full selected-tensor parity or explicitly qualified probe coverage; browser-tree memory and swap telemetry; experiment manifest and RUNS.md row

webgpu-ssb-preflight

MacBook Pro (M5 Max, 128 GB)

manual hardware diagnostic

parity preflight

PYTHONPATH=src python -m pytest -q tests/contracts/test_ssb_backend_contract.py tests/infrastructure/test_webgpu_sources.py tests/e2e/test_webgpu_widget_sync.py

pytest log; physical browser phase/loss parity and timing artifact