Kernel and benchmark dashboard#

This is the one-page technical overview of quantem.gpu: what the scientific kernels compute, where each runtime implements them, how parity is proved, and what the latest retained measurements actually mean.

Read the state before comparing the number

First-process source load, prepared-source load, warm resident compute, and saved-result reopen are different experiments. A binned or cropped source is never presented as native resolution. Check the complete provenance ledger before using a number in a design or release decision.

Dashboard review: 2026-08-22. Every current measured timing below shows the device, test date, and exact source revision. The overview omits opaque evidence IDs and does not replace the complete benchmark provenance ledger.

Coverage and next runs#

The filterable coverage registry now keeps every required configuration visible, including configurations that have never run. It separates complete measurements, partial evidence, pending work, refuted experiments, and fail-closed unsupported paths. Each open row names a stable runbook, its physical owner, and the exact artifact required for promotion.

State

Gate count

✓ Measured

18

◐ Partial

34

○ Pending

73

! Blocked

35

× Refuted

4

Not supported

27

↺ Superseded

0

Platform

Tracked gates

CUDA

17

Python MPS

37

Native Swift/Metal

56

WebGPU

58

CPU reference

23

Platform and computer coverage#

Each row identifies one reproducible hardware configuration. Counts describe tracked cells, including explicit unsupported contracts; a pending value remains a test to run. Load, admission, memory, and performance gates are multiplied across compatible computers because hardware changes the result. Platform-wide correctness or unsupported contracts are recorded once instead of creating misleading duplicate hardware rows.

Platform

Computer

Tracked cells

Measured

Partial

Pending

Blocked

Refuted

Unsupported

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

17

2

5

5

0

0

5

Python MPS

MacBook Air (M2, 8 GB)

9

0

0

7

2

0

0

Python MPS

MacBook Pro (M5 Max, 128 GB)

18

2

3

6

0

2

5

Python MPS

MacBook Pro (M5, 24 GB)

10

0

1

8

0

1

0

Native Swift/Metal

MacBook Air (M2, 8 GB)

14

0

0

11

2

0

1

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

27

8

5

7

0

0

7

Native Swift/Metal

MacBook Pro (M5, 24 GB)

15

4

3

7

0

0

1

WebGPU

MacBook Air (M2, 8 GB)

15

0

0

4

11

0

0

WebGPU

MacBook Pro (M5 Max, 128 GB)

27

2

4

7

9

1

4

WebGPU

MacBook Pro (M5, 24 GB)

16

0

0

7

9

0

0

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

2

0

0

2

0

0

0

CPU reference

MacBook Air (M2, 8 GB)

2

0

0

0

2

0

0

CPU reference

MacBook Pro (M5 Max, 128 GB)

2

0

2

0

0

0

0

CPU reference

MacBook Pro (M5, 24 GB)

2

0

0

2

0

0

0

CPU reference

Portable CI runner

15

0

11

0

0

0

4

Agents and maintainers should begin with:

python scripts/benchmark_registry.py next --limit 10
python scripts/benchmark_registry.py command GATE_ID

The overview tables below retain current measurements, qualified probes, and clearly marked diagnostics. Read the State cell before the time. The coverage registry is the canonical place to find missing combinations and their reproduction entry points.

Speed and memory at a glance#

The rows below are deliberately not a leaderboard: results with different fixtures, cache states, scientific plans, or wall-clock boundaries are not ranked against one another. Each row keeps those conditions together. Measurement tables are keyed first by Platform, then by the reproducible Computer class; local host nicknames are never public benchmark identifiers.

Current qualified load measurements#

One row is one exact configuration on one reproducible computer. Platform and computer are the first two columns; scan geometry, detector geometry, bin, dtype, cache state, statistic, memory observations, device, date, and revision remain separate data. The table is generated from the benchmark registry so a superseded result cannot remain the dashboard headline.

Only explicitly designated current rows appear here. Historical, superseded, refuted, and unmeasured configurations remain in the complete tables below rather than being silently deleted.

Platform

Computer

State

Selected scan

Source detector

Detector bin

Output detector

Source dtype

Staging dtype

Resident dtype

Scientific gate

Cache/process state

Wall boundary

Samples

p50

p95

Maximum

Logical resident

Accelerator/driver peak

Process/tree peak

Process physical-footprint peak

Swap delta

Parity

Device tested

Date tested

Revision

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

one same-process warmup; operating-system source pages uncontrolled; no eviction performed; fresh returned destination released after each trial

public io.load return after MPS synchronization; full-volume hash and release excluded

7

0.406624 s

0.428164 s

0.428164 s

18.000 GiB

18.442 GiB

18.692 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

uint8

uint16

exact

fresh QH5 index root and destination; macOS F_NOCACHE on source hashing and indexed descriptors; immutable source already audited

catalog, pipeline compilation, plan, complete exact private-resident volume, seven products, metadata, and provenance

7

0.577793 s

0.900979 s

0.900979 s

18.000 GiB

18.571 GiB

0.874 GiB

n/a

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

c0ea44465e6346a8436a0b74f491a04af0b5dc32

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.225984 s

0.231192 s

0.231192 s

4.500 GiB

5.087 GiB

0.657 GiB

5.336 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

6

0.200036 s

0.204694 s

0.204694 s

1.125 GiB

1.712 GiB

1.039 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

n/a

uint16

exact

original compressed source reread every visit; existing QH5 index and source-bound packing layout; exact DPC sums reused after first load; OS pages uncontrolled

indexed source open through synchronous complete packed-resident return; catalog and independent full-count audit separate

7

1.496660 s

2.086119 s

2.086119 s

18.000 GiB

n/a

n/a

2.531 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-09-06

e305f9216ed359397e27f69882931ffe16de8d99

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.678703 s

0.697179 s

0.697179 s

4.500 GiB

5.087 GiB

0.662 GiB

5.271 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

n/a

uint16

n/a

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.640942 s

0.651931 s

0.651931 s

1.125 GiB

1.712 GiB

1.018 GiB

2.319 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared immutable block indexes; explicitly warm source pages; fresh browser target per retained run; 699 pageouts; zero swap growth; no dropped runs

navigation through scientifically usable exact resident output and diagnostic frame checksums; exhaustive full-volume hash and application E2E excluded

7

1.358000 s

1.594000 s

1.594000 s

18.000 GiB

n/a

6.500 GiB

n/a

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

uint16

uint16

exact

prepared indexes; source pages unspecified; CPU exact-reference validation only

public CPU reference load; full-volume hash excluded

1

32.788039 s

32.788039 s

32.788039 s

18.000 GiB

n/a

19.141 GiB

n/a

0 B

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-22

68dbe3aa5f816e2d1c1ae976e1874790cffb4319

These measurements use a complete 512x512 scan and native 192x192 uint16 source detector, with scan bin 1 and no crop. The current Python MPS, Native Swift/Metal, and CPU rows use fixture real-512x512x192x192-u16-bslz4-27shard-master-fixture-c. The physical WebGPU smoke uses the separately fingerprinted full-native-webgpu-512x512x192x192-u16 fixture; it is not a cross-fixture speed comparison. Every row retains its own fixture SHA-256 and source identity in the complete registry. Detector bin is explicit. Exact binned uint16 output is source-identity-bound to its retained maximum-count audit; it is not a general license to narrow arbitrary input.

For this immutable source, the complete maximum-count audit of 53 proves that an 8x8 exact detector sum is at most 3,392, so bins through 8 fit in uint16. Other sources must pass their own complete range audit or use a provably sufficient wider integer dtype.

The package rows are prepared-index measurements. Python MPS source pages were uncontrolled after one same-process warmup. Native Swift/Metal uses prepared immutable indexes and controlled F_NOCACHE source descriptors. Neither is a cold arbitrary-source or application end-to-end result. Historical CUDA, WebGPU, earlier CPU, destination-reuse, and superseded measurements remain in Verified benchmark results and the complete coverage registry; they are not silently discarded or mixed into this current table.

Logical resident bytes, accelerator/driver peak, process RSS high-water, process physical-footprint peak, and swap delta are distinct observations and are not additive. A missing physical-footprint value is displayed as n/a, not inferred from RSS. Process RSS does not include every direct Metal allocation, so it cannot replace accelerator or physical-footprint telemetry. For example, the current full-native Python MPS row reports 18.442 GiB of driver allocation for an 18.00 GiB logical resident tensor; that is a measured after-load boundary, not a universal peak estimate. On the 24 GB MacBook Pro, the full-native bin-1 smoke passed every exact volume and product hash but produced approximately 723.31 MiB additional swap use and 758,448,128 B of swapouts. The safety gate therefore stopped after one trial; it is partial evidence, not a repeatable performance distribution.

Dtype support and peak memory#

“Source,” “staging,” “accumulation,” and “resident” dtype describe different stages. A uint8 row is scientifically exact only when the source is already uint8 or a source-identity-bound complete audit proves maximum <= 255 and pixelsAbove255 == 0. Otherwise a saturating uint8 output, such as the WebGPU clip8 browse decode, clips values above 255 and is a browse representation, not raw-count evidence.

The registry below separates those cases. Each row fixes one platform, computer class, scan, detector, bin, crop, source dtype, staging dtype, resident dtype, and scientific gate. Browse-only can be benchmarked, but it cannot satisfy an exact scientific gate.

Each row is one atomic source, staging, and resident dtype contract. Exact rows preserve scientific counts; browse-only rows are explicit saturating representations and cannot satisfy an exact gate.

Platform

Computer

State

Selected scan

Source detector

Scan bin

Detector bin

Output detector

Crop

Source dtype

Staging dtype

Resident dtype

Scientific gate

Precision contract

Implementation basis

Next gate or reason

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/cuda/kernels/bslz4.cu shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=cuda) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained.

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public CUDA io.load and retain its complete-volume hash, source/resident provenance and peak memory.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The CUDA decoder has a dedicated uint16 bitshuffle path and existing unit parity; the current registry has no uncontended full-volume native uint16 distribution.

Run the public CUDA load entry point on an uncontended device and retain a complete-volume uint16 hash plus source/staging/resident provenance and peak allocation.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form.

CUDA io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output.

CUDA io.load has no uint8 output dtype.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output.

CUDA io.load has no saturating uint8 browse output.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening.

CUDA io.load has no uint16-to-uint32 widening.

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for CUDA.

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/mps/kernels/bslz4.msl shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=mps) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained.

Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public Python MPS io.load and retain its complete-volume hash, source/resident provenance and peak memory.

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The MPS uint16 decoder and retained full-volume canonical hash preserve native counts at detector bin 1.

n/a

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form.

Python MPS io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output.

Python MPS io.load has no uint8 output dtype.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output.

Python MPS io.load has no saturating uint8 browse output.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening.

Python MPS io.load has no uint16-to-uint32 widening.

Python MPS

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for Python MPS.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

Native4DSTEMIO can describe one-byte indexed sources, but Metal4DSTEMIndexedLoader.swift requires a uint16 source and uint16 staging at the public integrated boundary.

The integrated indexed Swift/Metal loader does not admit a native uint8 source.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

Metal4DSTEMIndexedLoader.swift supports the uint16 fallback stage, while retained optimized measurements used separately registered audited uint8 staging.

Run a source whose complete audit does not authorize uint8 staging and retain exact uint16 staging-to-resident parity and allocation evidence.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

NativeHDF5Bridge.swift describes one- and two-byte indexed sources, and Metal4DSTEMIndexedLoader.swift requires uint16 input.

The native indexed Swift/Metal source contract does not admit uint32 HDF5 detector values.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

Metal4DSTEMExactBinner.provenance accepts only uint16 or uint32 output, although it can use an audited uint8 staging buffer.

The exact native Swift/Metal load contract cannot publish a uint8 resident scientific volume.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

The indexed exact loader exposes uint16 or uint32 scientific output and has no separate saturating browse-resident API.

Explicit saturating uint8 resident output is not implemented by the public native Swift/Metal load boundary.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

Metal4DSTEMExactBinner has tested uint16-to-uint32 kernels and provenance, but Metal4DSTEMIndexedBinnedLoad fixes integrated resident output to uint16.

Expose uint32 output through the public indexed loader and cache contract, then retain complete load, metadata, and physical parity.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

Native4DSTEMValueRangeAudit binds source identity and maximum counts; the retained controlled full-native run used audited uint8 staging and reproduced exact uint16 volume and product hashes.

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

h5reader.ts, bslz4.ts, and local-h5.ts recognize matching native uint8 source, decode, and resident modes; no retained full-volume physical uint8 fixture proves the complete path.

Run an exact physical hardware-browser full-volume uint8 source through local HDF5 decode and retain output hash, allocation, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The physical WebGPU path retained exact native uint16 resident checks and seven runs, but device allocation and timed complete-volume hashing remain incomplete.

Add WebGPU device-allocation telemetry and retain an exact complete-volume hash inside the timed boundary before promoting the integrated load gate.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

h5reader.ts, bslz4.ts, and local-h5.ts implement matching native uint32 source/decode/resident modes; physical full-volume proof is absent.

Run a native uint32 HDF5 source on physical hardware WebGPU and retain the complete resident hash, browser/device memory, and source/decode/resident provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

bslz4.ts has an audit-dependent low8 kernel, but local-h5.ts does not bind a complete audit identity to the returned uint8 resident state.

Replace the experimental low8 global flag with a typed source-identity-bound audit in the local-HDF5 public result, then retain exact full-volume parity.

WebGPU

MacBook Pro (M5 Max, 128 GB)

○ Pending

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

bslz4.ts decodes all uint16 planes and saturates to 255; source tests distinguish this from the experimental low8 audit path.

Run the fused full-plane clip8 path on physical hardware with values above 255 and retain output hash, browser/device memory, and browse-only provenance.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

local-h5.ts requires decode dtype uint32 to match a uint32 HDF5 source and rejects a uint16 source request.

The WebGPU local-HDF5 public path does not widen uint16 source values into uint32 resident storage.

WebGPU

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

local-h5.ts keeps decode and resident integer mode matched for native paths and rejects mismatched uint32 requests; it exposes no uint8-stage-to-uint16 reconstruction contract.

Audited uint8 staging followed by exact uint16 resident reconstruction is not implemented in the public WebGPU local-HDF5 path.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint8

uint8

uint8

exact

native integer counts

src/quantem/gpu/io/hdf5/cpu.py returns the HDF5 native dtype unchanged at detector bin 1; a complete public-API uint8 fixture artifact is not retained.

Add a repository fixture that proves public io.load preserves every native uint8 count, shape, order, and provenance field at detector bin 1.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint16

exact

native integer counts

The CPU reference returns native uint16 counts at detector bin 1 and has retained physical reference probes; the portable registry gate is not a complete public-API artifact.

Promote only after the portable public API fixture retains complete-volume uint16 hash, shape, order, mask, and metadata parity.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

1

192 × 192

none

uint32

uint32

uint32

exact

native integer counts

The CPU reference preserves the HDF5 native dtype at detector bin 1; public load defaults may auto narrow uint32, so the exact uint32 request must be explicit and tested.

Add a public-API native uint32 fixture that disables advisory auto narrowing and proves complete count, dtype, shape, and metadata parity.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint8

exact

source-identity-bound complete value-range audit

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint8 output.

CPU reference io.load has no uint8 output dtype.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint8

browse-only

explicit saturation to 255

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no saturating uint8 browse output.

CPU reference io.load has no saturating uint8 browse output.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint16

uint32

exact

lossless integer widening

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint32 widening.

CPU reference io.load has no uint16-to-uint32 widening.

CPU reference

Portable CI runner

Not supported

512 × 512

192 × 192

1

1

192 × 192

none

uint16

uint8

uint16

exact

source-identity-bound uint8 staging with exact uint16 reconstruction

src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no audited uint8 staging.

Audited uint8 staging with uint16 resident reconstruction is not implemented for CPU reference.

This capability matrix does not infer peak memory. The current load table is authoritative only for its matching shape, dtype, device, cache state, and wall boundary. In particular, browser RSS may be smaller than a WebGPU resident payload because it does not capture every device allocation; that is an incomplete peak, not evidence that the payload disappeared.

The current C matrix measures native uint16 input. The full-native logical payload is exactly 19,327,352,832 bytes (18.00 GiB). Python io.load on CUDA and MPS keeps that complete native detector and holds it ANS encoded, so its resident bytes depend on the acquisition, not on a detector bin. Detector bins are explicit options of the native Swift/Metal and WebGPU loaders and exact views of the encoded resident in the CUDA browse service. The native Swift/Metal primitive has admitted one full 18.00 GiB resident load on the 24 GB Apple M5, but repeated uncontended memory and paging evidence remains incomplete. A consumer may apply a conservative working-set limit; that application policy is separate from primitive capability and may not silently select detector bin 2.

The current production WebGPU detector-bin-2/4/8 path accumulates into and stores float32. Those historical timings remain useful implementation history, but they do not satisfy the exact-integer resident contract. The platform/computer registry therefore marks exact WebGPU bins 2/4/8 blocked until integer accumulation and residency pass full-volume parity. Seven-run physical exact full-volume distribution evidence remains in the retained-measurement table; unavailable WebGPU device-allocation telemetry is still shown as n/a, never inferred from browser RSS.

Public Python io.load keeps stored counts: omit dtype or pass dtype="native". Its one conversion is an explicit dtype="scaled_uint16" for float32 sources, which stores calibrated integer codes, stays ANS encoded, and reports the measured conversion error. The CUDA, Python MPS and CPU reference rows above for uint8 outputs, uint32 widening and audited uint8 staging are therefore Not supported: those were load-time casts that io.load no longer offers.

Minimum-device memory gates#

The public release floors are 6 GiB of dedicated VRAM for CUDA and 8 GB of total laptop RAM for WebGPU. The WebGPU number is the entire machine budget shared by the operating system, browser, JavaScript heap, staging buffers, and GPU—not memory available exclusively to one GPUBuffer.

✓ is awarded only after the complete load-and-product pipeline runs on a physical device at or below the stated floor with retained peak memory, pressure/swap, output parity, and responsiveness evidence. A calculated payload can prove No, but it can only establish a Pending candidate. An allocator cap on a larger GPU is a useful pre-check, not physical-device signoff.

Platform

Computer

Minimum device

Selected scan

Scan plan

Source detector

Detector bin

Output detector

Resident dtype

Resident payload

Gate

Reason

Device tested

Date tested

CUDA

Physical 6 GiB CUDA computer pending

6 GiB VRAM

512x512

Full

192x192

1

192x192

uint16, ANS encoded

1.98 to 3.14 GiB

Pending

Measured encoded payload of three real acquisitions; complete physical 6 GiB peak not retained

—

—

Python MPS

MacBook Air (M2, 8 GB)

8 GB unified RAM

512x512

Full

192x192

1

192x192

uint16, ANS encoded

Data-dependent

Pending

Encoded payload and physical 8 GB pressure/parity run not retained

—

—

Native Swift/Metal

MacBook Air (M2, 8 GB)

8 GB unified RAM

512x512

Full

192x192

1

192x192

uint16

18.00 GiB

Blocked

Resident payload exceeds physical memory

—

—

Native Swift/Metal

MacBook Air (M2, 8 GB)

8 GB unified RAM

512x512

Full

192x192

2

96x96

uint16

4.50 GiB

Pending

Current clean-revision pressure/parity run not retained

—

—

Native Swift/Metal

MacBook Air (M2, 8 GB)

8 GB unified RAM

512x512

Full

192x192

4

48x48

uint16

1.125 GiB

Historical

Earlier physical run passed; current clean-revision repeat remains pending

Apple M2 MacBook Air (Mac14,2, 8 GB)

2026-08-18

Native Swift/Metal

MacBook Air (M2, 8 GB)

8 GB unified RAM

512x512

Full

192x192

8

24x24

—

—

Not supported

Current native load-plan contract supports bins 1, 2, and 4

—

—

WebGPU

MacBook Air (M2, 8 GB)

8 GB total RAM

512x512

Full

192x192

1

192x192

uint16

18.00 GiB

Blocked

Resident payload exceeds total machine RAM

—

—

WebGPU

MacBook Air (M2, 8 GB)

8 GB total RAM

512x512

Full

192x192

2

96x96

uint16

4.50 GiB

Blocked

Production detector binning stores float32; exact-integer accumulation and residency must be integrated before the physical pressure/parity gate

—

—

WebGPU

MacBook Air (M2, 8 GB)

8 GB total RAM

512x512

Full

192x192

4

48x48

uint16

1.125 GiB

Blocked

Production detector binning stores float32; exact-integer accumulation and residency must be integrated before the physical pressure/parity gate

—

—

WebGPU

MacBook Air (M2, 8 GB)

8 GB total RAM

512x512

Full

192x192

8

24x24

uint16

0.28125 GiB

Blocked

Production detector binning stores float32; exact-integer accumulation and residency must be integrated before the physical pressure/parity gate

—

—

Python io.load on CUDA and MPS has no detector bin: it keeps the complete native detector ANS encoded, so its payload depends on the counts. Three real 512x512x192x192 uint16 acquisitions encode to 2,127,983,353, 2,279,376,159 and 3,374,307,987 bytes (0.11 to 0.18 of the 18.00 GiB of counts) on the CUDA loader (NVIDIA RTX PRO 6000 Blackwell, 2026-10-05). The exact integer WebGPU bin-2/4/8 payloads can fit within the nominal 8 GB total-RAM floor, but payload arithmetic alone is not acceptance. Production WebGPU still needs integer accumulation/residency plus a physical browser run that captures browser, adapter, staging, operating-system, pressure, swap, and scientific parity. Likewise, the current Blackwell timings do not prove a physical 6 GiB CUDA floor.

What a 4 or 6 GiB budget can hold#

This capacity chart fixes the full scan at 512x512 and the native detector at 192x192. “Payload” excludes decoder scratch, staging buffers, allocator reserve, and other GPU users unless the row reports a measured process peak.

Detector bin

Resident dtype

Resident payload

4 GiB fit

6 GiB fit

Evidence state

1

uint16

18.00 GiB

No

No

Calculated dense payload (native Swift/Metal and WebGPU bin 1)

1

uint16, ANS encoded

1.98 to 3.14 GiB

Candidate

Candidate

Measured CUDA io.load payload of three real acquisitions; physical 4/6 GiB signoff Pending

1

uint8

9.00 GiB

No

No

Complete-audit lossless path only (WebGPU low8)

4

uint16

1.125 GiB

Candidate

Candidate

Physical 8 GB M2 Air evidence (native Swift/Metal); 4/6 GiB signoff Pending

The encoded CUDA/MPS load needs no memory plan of its own: io.load reads and encodes the acquisition in bounded scan blocks, and read(...) and the detector kernels decode bounded blocks, so the working set is the encoded payload plus bounded decode scratch.

Each 512x512 float32 product map is only 1 MiB, and one 192x192 float32 mean diffraction pattern is 144 KiB. The source working set—not the final BF/ADF/DF/DPC image—is the capacity problem.

For a full scan with output detector bin \(b\) and \(w\) resident bytes per value, the payload alone is

\[ B_{\mathrm{payload}} =N_{R_r}N_{R_c} \left\lceil\frac{N_{k_r}}{b}\right\rceil \left\lceil\frac{N_{k_c}}{b}\right\rceil w. \]

Peak memory is larger: the benchmark must also report live compressed bytes, decode and reduction scratch, staging/upload buffers, allocator reserve, products, concurrent GPU users, and—on unified memory—process pressure and swap. A calculated payload is never relabeled as a measured peak.

Small-GPU support today

The CUDA/MPS screening.prepare path loads the complete acquisition once into encoded storage and emits mean DP, total, BF, ABF, ADF, DF, CoM, rotation, and iDPC without cropping the scan. Physical 4 and 6 GiB product-pipeline signoff is Pending. See the screening API before choosing a plan.

Platform-first module dashboard#

Every table keeps the scientific module as its section and puts the execution platform in the first column. Empty cells are forbidden:

  • ✓ — verified with retained real-data or physical-device parity evidence.

  • Test — deterministic source/test coverage without an equivalent retained physical or native-data run.

  • Pending — implementation exists, but that exact size, bin, or timing evidence has not been retained yet.

  • Ref — CPU correctness adjudication, never a production fallback.

  • — — unsupported or not a target.

I/O and first usable product — quantem.gpu.io#

Exact configuration gaps#

The measured table above owns retained full-scan timings. The filterable coverage registry owns every missing combination as one Platform → Computer → detector-bin → cache-state row; this page does not maintain a second bundled list. Real-space crop remains an explicit selector and is never an implicit performance policy.

Selective scan rectangles#

The physical WebGPU rows below are prepared whole-shard-selective rectangle loads. A retained frame-span manifest lets the loader omit nonintersecting shards and decode/upload only selected scan rows. It does not issue byte-range reads inside an intersecting shard.

Platform

Computer

Selected scan

Rectangle (row_start,row_stop,column_start,column_stop)

Source detector

Source dtype

Shards read

Storage bytes read

Samples

Loader p50

Loader p95

Loader maximum

Logical resident

Browser-tree RSS peak

Observed swap delta

Parity

Device tested

Date tested

WebGPU

MacBook Pro (M5 Max, 128 GB)

64x64

(0,64,0,64)

192x192

uint16

4 of 27

488,224,242 B

5

0.147 s

0.1544 s

0.156 s

301,989,888 B

1,724,317,696 B

0 B

✓

Chrome 151, Apple M5 Max Metal

2026-08-22

WebGPU

MacBook Pro (M5 Max, 128 GB)

256x256

(128,384,128,384)

192x192

uint16

14 of 27

1,705,556,941 B

5

0.381 s

0.3924 s

0.394 s

4,831,838,208 B

3,002,875,904 B

0 B

✓

Chrome 151, Apple M5 Max Metal

2026-08-22

WebGPU

MacBook Pro (M5 Max, 128 GB)

384x384

(64,448,64,448)

192x192

uint16

20 of 27

2,432,636,897 B

5

0.574 s

0.582 s

0.584 s

10,871,635,968 B

3,896,934,400 B

0 B

✓

Chrome 151, Apple M5 Max Metal

2026-08-22

All rows use prepared block indexes and a prepared frame-span manifest. The operating-system source-page state was uncontrolled/unspecified and no eviction was performed. Frame-manifest read/encoding, DevTools injection, checksum harness, products, and application E2E are excluded. A negative control without the frame-span manifest read all 27 shards (3.17 GB), so it is explicitly not selective evidence. CUDA and Python MPS load the complete acquisition and return a rectangle with read(scan_region=..., detector_region=...) from the encoded resident; native Swift/Metal bulk selection and physical cross-backend timing remain pending.

The parity check mark means all five runs for that rectangle passed independent CPU uint16 checksum probes at the first, middle, and last retained raw frames, plus output shape, dtype, row-major order, and selection metadata. Full-tensor readback parity was not performed. The WebGPU fixture view records source identity 1be810b9...; fixture C records c9c0d968..., and the CUDA master record is 4802ec16.... These rows are not a cross-lane fixture-controlled comparison.

Screening and prepared-product caches — quantem.gpu.screening#

Platform

Computer

Support

Operation

Source plan

Statistic

Time

Device tested

Date tested

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

✓

Exact streamed screening

Full 512x512x192x192 uint16; no crop/bin; warm source pages unspecified; empty result cache

p50 of 6

1.205 s

NVIDIA RTX PRO 6000 Blackwell Workstation Edition, GPU 0

2026-08-22

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓

Exact screening build

Full 512x512x192x192 uint16; no crop/bin; exact fallback pass

Single run

6.711 s

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

✓

Validated screening-v3 reopen

Prepared derived products

p50

20.803 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Native Swift/Metal

MacBook Air (M2, 8 GB)

✓

Validated exact-summary reopen

Prepared derived products

p50

0.029 s

Apple M2 MacBook Air (Mac14,2, 8 GB)

2026-08-19

WebGPU

N/A (not implemented)

—

Prepared screening workflow

—

—

—

—

—

CPU reference

Portable CI runner

Ref

Independent adjudication

Reference fixtures

—

Pending

—

—

The promoted MPS build derives its detector mean and masks from the full scan. When the provisional and final masks differ, BF/DF are recomputed exactly. The first-chunk-mask candidate failed parity and is retained only as a rejected experiment; its timing is not repeated on a current page. Saved-product reopen is likewise never presented as source load.

The current CUDA pinned-slot candidate reduced the like-for-like package p50 from 1.356516 s to 1.204713 s; candidate p95/max were 1.325760/1.329731 s. All six public arrays were byte exact in every trial. This is warm/source-pages-unspecified exact screening, not cold HDF5 and not a prepared-result reopen. Detailed pinned-registration and memory measurements remain in the provenance ledger.

Virtual images — quantem.gpu.detector#

Complete platform, computer, and detector-bin matrix#

The registry keeps one row per exact product configuration instead of hiding an unrun combination. Use the Platform, Computer, State, and Detector bin filters to isolate the implementation or physical computer you care about. A pending, blocked, or unsupported row is intentionally visible; only a measured row publishes a timing distribution.

The 8 GB MacBook Air detector-bin-2 row is the locked downstream acceptance configuration. Detector bins 4 and 8 remain separate library-capability cells, not silent application fallbacks. Detector bin 1 is explicitly blocked where its 18 GiB exact resident input cannot be admitted.

Each row fixes one platform, reproducible computer, detector bin, and exact product-suite boundary. Partial diagnostics retain their sample and memory context, but only fully measured gates display p50, p95, and maximum timing.

Platform

Computer

State

Selected scan

Source detector

Detector bin

Output detector

Resident dtype

Cache/process state

Wall boundary

Samples

p50

p95

Maximum

Logical resident

Accelerator/driver peak

Process/tree peak

Swap delta

Parity

Device tested

Date tested

Revision

Next gate or reason

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

warm operating-system source pages; unspecified; empty screening-result cache

quantem.gpu.screening.prepare package wall through exact complete six-array screening result in balanced A-B-B-A-B-A-A-B comparison

6

n/a

n/a

n/a

n/a

2.084 GiB

1.958 GiB

0 B

Pass

NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03

2026-08-22

023a6c497b106b216c87205d3fbec63377d77177

Run the current public ScreeningResult on an uncontended CUDA device with benchmark_screening.py –require-exact-full-suite; retain all nine hashes, uint64 total/ABF/ADF dtype and shape, p50/p95/maximum, and complete memory telemetry.

Python MPS

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before decoder, allocator, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming product path that never requires an 18 GiB resident input before attempting this physical row.

Python MPS

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

exact resident detector-bin-1 MPS input; source pages uncontrolled or warm; no prepared product result; five pageouts; zero swapout growth

single first-execution diagnostic only; synchronized detector-product publication was 0.499909 seconds; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

27.510 GiB

20.436 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

b8df61b55920cb098e2848301f6d30c45870511d

Repeat the exact product suite without pageout growth for synchronized p50, p95, and maximum timing plus complete memory telemetry.

Python MPS

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native MPS input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run only after an uncontended pressure preflight; use a single exact full-native smoke with a hard stop on swap/pageout growth before considering repeated product timing.

Native Swift/Metal

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB private-resident input exceeds the computer’s 8 GB unified memory before decoder, Metal, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming or mapped product path that never requires an 18 GiB private-resident input before attempting this physical row.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 Metal input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the complete product suite from the explicit exact detector-bin-2 resident volume with full parity, synchronized timing, pressure, swap, and peak memory.

Native Swift/Metal

MacBook Air (M2, 8 GB)

○ Pending

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 Metal input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the complete exact detector-bin-4 product suite on the current clean revision with full hashes, synchronized timing, pressure, compressed memory, swap, and peak memory; detector bin 2 remains the locked application acceptance policy.

Native Swift/Metal

MacBook Air (M2, 8 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

prepared immutable QH5 indexes; source pages unspecified

complete exact resident volume, products, metadata, calibration, and provenance; full private readback/hash excluded

6

0.313870 s

0.318865 s

0.318865 s

18.000 GiB

18.571 GiB

0.639 GiB

n/a

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

531e5001f5a0a886c058f772ef42d770f000890b

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.225984 s

0.231192 s

0.231192 s

4.500 GiB

5.087 GiB

0.657 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded

6

0.200036 s

0.204694 s

0.204694 s

1.125 GiB

1.712 GiB

1.039 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; 128 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified

package exact indexed load plus exact full-scan products; app UI, index creation, and full-volume hash excluded

1

n/a

n/a

n/a

18.000 GiB

18.571 GiB

0.624 GiB

n/a

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

0b305f69fc1039933463986a0f74609e22d4dd35

Repeat uncontended full-native publication with complete product hashes, p50/p95/max, driver allocation, RSS, pressure, and swap.

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

2

96 × 96

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.678703 s

0.697179 s

0.697179 s

4.500 GiB

5.087 GiB

0.662 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

✓ Measured

512 × 512

192 × 192

4

48 × 48

uint16

prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials

exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded

6

0.640942 s

0.651931 s

0.651931 s

1.125 GiB

1.712 GiB

1.018 GiB

0 B

Pass

Apple M5 10-core integrated GPU; 24 GB unified memory

2026-08-22

5106ca48349231549c440b40609983fd3a8dacda

n/a

Native Swift/Metal

MacBook Pro (M5, 24 GB)

Not supported

512 × 512

192 × 192

8

24 × 24

uint16

not applicable

unsupported load-plan contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not applicable

n/a

n/a

n/a

Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked exact-residency contract

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB total memory before browser, adapter, decoder, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming browser product path that never requires an 18 GiB resident input before attempting this physical row.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

✓ Measured

512 × 512

192 × 192

1

192 × 192

uint16

prepared QH5 indexes; warm source pages; exact full-native uint16 resident input; fresh browser target per retained run

complete internal GPU product suite from exact resident input through awaited readbacks; parity hashing, sorting, and encoding excluded

7

0.482000 s

0.483600 s

0.483600 s

18.000 GiB

n/a

6.492 GiB

0 B

Pass

Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false

2026-08-22

d8e6f562a8ab43086a4ddea9eecfe6fd26b7beea

n/a

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5 Max, 128 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native hardware-WebGPU input

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run a guarded physical hardware-browser admission smoke, then retain the complete native product suite with full parity, timing, browser/device memory, pressure, and swap if admission remains safe.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

2

96 × 96

uint16

resident exact detector-bin-2 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

4

48 × 48

uint16

resident exact detector-bin-4 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

WebGPU

MacBook Pro (M5, 24 GB)

! Blocked

512 × 512

192 × 192

8

24 × 24

uint16

resident exact detector-bin-8 hardware-WebGPU input required

complete product suite publication

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory.

CPU reference

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native CPU input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run the independent full-native CPU reference product suite only when host CPU, storage, and memory ownership are uncontended; retain all hashes, timing, and peak RSS.

CPU reference

MacBook Air (M2, 8 GB)

! Blocked

512 × 512

192 × 192

1

192 × 192

uint16

not admissible

blocked reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Not run

n/a

n/a

n/a

The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before products and operating-system headroom. Next: Implement and prove an exact bounded-streaming CPU reference product path that never requires an 18 GiB resident input before attempting this physical row.

CPU reference

MacBook Pro (M5 Max, 128 GB)

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

exact full-native CPU-resident input after a warm operating-system page-cache source load

single independent detector-product traversal was 31.078040 seconds; source load excluded; no p50, p95, or maximum distribution

1

n/a

n/a

n/a

18.000 GiB

n/a

36.450 GiB

n/a

Pass

Apple M5 Max CPU; 128 GB unified memory

2026-08-19

334b7b5135fe29787540370a00f280fa138430a2

Add ABF, retain all product hashes, and repeat the exact physical CPU reference traversal for p50/p95/maximum and complete process-memory telemetry.

CPU reference

MacBook Pro (M5, 24 GB)

○ Pending

512 × 512

192 × 192

1

192 × 192

uint16

resident exact full-native CPU input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Run only after an uncontended memory preflight; retain one exact full-native CPU reference smoke, complete product hashes, pressure, swap, and peak RSS before repetition.

CPU reference

Portable CI runner

◐ Partial

512 × 512

192 × 192

1

192 × 192

uint16

resident reference input

complete reference product suite

n/a

n/a

n/a

n/a

n/a

n/a

n/a

n/a

Pending

n/a

n/a

n/a

Keep exact reference tests current; no CPU speed claim is required.

Retained operation-level timings#

Platform

Computer

Operation

Scan grid

Detector

Detector bin

Input state

Statistic

Time

Device tested

Date tested

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Mean diffraction

512x512

192x192

1

Warm resident

p50

18.392 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

BF exact sum

512x512

192x192

1

Warm resident

p50

3.768 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

ADF exact sum

512x512

192x192

1

Warm resident

p50

5.586 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

DF exact sum

512x512

192x192

1

Warm resident

p50

3.747 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

Mean diffraction

512x512

48x48

4

Warm resident

p50

74.805 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

BF exact sum

512x512

48x48

4

Warm resident

p50

2.502 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

ADF exact sum

512x512

48x48

4

Warm resident

p50

4.404 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

DF exact sum

512x512

48x48

4

Warm resident

p50

2.642 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Fused BF, ABF, ADF, total, and row/column moments

512x512

192x192

1

Controlled source load

p50

119.040 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-22

Native Swift/Metal

MacBook Air (M2, 8 GB)

BF, ABF, ADF, total, and row/column moments

512x512

48x48

4

Prepared resident-cache fallback

Single run

103.0 ms

Apple M2 MacBook Air (Mac14,2, 8 GB)

2026-08-19

Native Swift/Metal

MacBook Air (M2, 8 GB)

Exact moment widening for resident summary

512x512

48x48

4

Same validated fused source pass; no resident traversal

Single run

0.569 ms

Apple M2 MacBook Air (Mac14,2, 8 GB)

2026-08-20

WebGPU

MacBook Pro (M5 Max, 128 GB)

Mean diffraction

512x512

192x192

1

Warm resident

p50

50.9 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

BF exact sum

512x512

192x192

1

Warm resident

p50

5.5 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

ADF exact sum

512x512

192x192

1

Warm resident

p50

15.0 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

DF exact sum

512x512

192x192

1

Warm resident

p50

43.4 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

CPU reference

MacBook Pro (M5 Max, 128 GB)

Virtual-image adjudication

512x512

192x192

1

One independent traversal

Single run

31.08 s

Apple M5 Max CPU

2026-08-19

The current integer and mean-DP rows pass their independent CPU reference. CUDA uses native detector resolution on fixture D; MPS uses explicit detector bin 4 on fixture C. WebGPU uses native detector resolution on D. The current native bin-1 source pass produces three virtual images, detector sums, and exact uint32 total and detector moments in one fused stage. That stage overlaps source IO and is not additive with the package-wall measurement. Revision 70bc366 proves when bin-4 accumulators fit, then widens the three moment maps to uint64 in one small dispatch. The 0.569 ms row is only that incremental widening; it is not the full source pass. The cache-only fallback row remains because it has a different input state. Neither historical bin-4 row is a compressed-source load time.

Detector moments and phase contrast — quantem.gpu.dpc#

Platform

Computer

Operation

Scan grid

Detector

Detector bin

Input state

Statistic

Time

Device tested

Date tested

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

CoM row and column

512x512

192x192

1

Warm resident

p50

13.002 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Fixed-orientation iDPC

512x512

192x192

1

CPU small-field integration

p50

21.272 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

CoM row and column

512x512

48x48

4

Warm resident

p50

4.637 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

Fixed-orientation iDPC

512x512

48x48

4

CPU small-field integration

p50

12.678 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Native Swift/Metal

MacBook Air (M2, 8 GB)

CoM, DPC, iDPC, and display statistics

512x512

48x48

4

Prepared exact uint64 moments

Single run

11.389 ms

Apple M2 MacBook Air (Mac14,2, 8 GB)

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

DPC row

512x512

192x192

1

Warm cached CoM

p50

0.9 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

DPC column

512x512

192x192

1

Warm cached CoM

p50

0.7 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

iDPC

512x512

192x192

1

Explicit 0-degree rotation

p50

1.4 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

CPU reference

MacBook Pro (M5 Max, 128 GB)

Rotation and iDPC adjudication

512x512

192x192

1

CPU reference

Single run

177.6 ms

Apple M5 Max CPU

2026-08-19

The current CUDA and MPS CoM rows pass their respective independent references. The native row is the complete derived-product stage after exact moments exist; it includes center/mean, alignment, Metal iDPC, and float-surface construction. It does not include resident-cache traversal or source loading. The CUDA/MPS cross-fixture timing rows are not compared numerically. The prior same-fixture detector-bin-4 comparison remains a historical block for iDPC at 2.84e-5 maximum error. On the retained M5 Max hardware run, WebGPU DPC row and column were byte exact. Optimized and zero-rotation iDPC had zero frozen tolerance violations; optimized rotation had maximum absolute error 1.52587890625e-5 and maximum tolerance ratio 0.8993483035.

Single-sideband ptychography — quantem.gpu.SSB#

These are square scan-grid sizes, not detector dimensions.

Current 512x512 operation timing is separated from the size-support matrix. Source loading, G(\mathbf k,\boldsymbol{\nu}) preparation, and UI paint are excluded.

Platform

Computer

State

Operation

Detector plan

BF policy

Boundary

Statistic

Time

Device tested

Date tested

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Partial

Complex object

Native 192x192

8,928 active

Warm resident GPU

p50

13.883 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Partial

Exact phase

Native 192x192

8,928 active

Warm resident GPU

p50

32.035 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Partial

Exact phase and loss

Native 192x192

8,928 active

Warm resident GPU

p50

32.335 ms

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

Partial

Exact phase and loss

Explicit detector bin 2 to 96x96

2,275 calibrated

Single synchronized reconstruction

Single run

497.187 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Partial

Complex object

Native-detector exact BF columns

9,074 logical / 2,459 executed

Warm complete Hermitian cache

p50

8.911 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Partial

Exact phase-variance loss

Native-detector exact BF columns

9,074 logical / 2,459 executed

Warm complete Hermitian cache

p50

25.120 ms

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

Refuted diagnostic

Complex object

Native 192x192 companion

3,418 active

Readback-complete compute wall

p50

32.5 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

Refuted diagnostic

Exact phase

Native 192x192 companion

3,418 active

Readback-complete compute wall

p50

102.1 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

WebGPU

MacBook Pro (M5 Max, 128 GB)

Refuted diagnostic

Exact phase and loss

Native 192x192 companion

3,418 active

Readback-complete compute wall

p50

189.4 ms

Chrome 151, Apple M5 Max Metal-3

2026-08-19

CPU reference

Portable CI runner

Reference

SSB

Frozen adjudication only

Frozen fixture

Reference

—

Pending

—

—

The WebGPU SSB values are retained diagnostic timings, not accepted scientific performance. All 262,144 phase values differed from the frozen reference and the wrapped maximum error was 0.0597773 rad against a 0.0002 rad gate. That implementation remains refuted until phase parity is restored.

The raw detector-bin-2 MPS phase agrees with an independent CUDA reference to 1.2815e-6 wrapped radians maximum; loss differs by 7.45e-9. A prepared BF-column companion from the same campaign is rejected because its stored columns do not match the declared detector-bin coordinate grid. Its faster timings are not published as scientific results.

Calibration is a separate operation:

Platform

Computer

Search

Refinement

Repetitions

Statistic

Time

Result

Device tested

Date tested

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

Seeded Optuna TPE, 200 trials

Nelder–Mead

3

p50

11.168 s

Byte-deterministic parameters, phase, object, and loss

NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1

2026-08-19

Python MPS

MacBook Pro (M5 Max, 128 GB)

Optuna TPE, 200 trials

Nelder–Mead

—

—

Pending

Current compatible source not profiled

—

—

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

Seeded TPE, 200 trials

Nelder–Mead

3

p50

6.061 s

Deterministic parameters and loss

Apple M5 Max (Mac17,6, 40-core GPU)

2026-08-19

WebGPU

N/A (not implemented)

—

—

—

—

—

Unsupported

—

—

CPU reference

Portable CI runner

—

—

—

—

—

Reference only

—

—

Levenberg–Marquardt is not implemented in any current SSB backend. An earlier CUDA atomic-objective calibration split into two different fitted minima under an identical seed; it is retained as a rejected experiment, not a benchmark row.

Platform

Computer

Scan grid

Source kind

BF policy

State

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

128x128

Fixed-size parity

Frozen fixture

Test

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

256x256

Fixed-size parity

Frozen fixture

Test

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

512x512

Native real acquisition

Full active BF

Partial physical evidence

CUDA

Linux CUDA workstation (dual 96 GB Blackwell GPUs)

1024x1024

Fixed-size parity

Frozen fixture

Test

Python MPS

MacBook Pro (M5 Max, 128 GB)

128x128

Resized/synthetic

Fixed-size fixture

Test

Python MPS

MacBook Pro (M5 Max, 128 GB)

256x256

Resized/synthetic

Fixed-size fixture

Test

Python MPS

MacBook Pro (M5 Max, 128 GB)

512x512

Native real acquisition

Full active BF

Partial physical evidence

Python MPS

MacBook Pro (M5 Max, 128 GB)

1024x1024

Synthetic

Fixed-size fixture

Test

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

128x128

—

—

Not supported

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

256x256

—

—

Not supported

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

512x512

Native real acquisition

Full active BF

Partial physical evidence

Native Swift/Metal

MacBook Pro (M5 Max, 128 GB)

1024x1024

—

—

Not supported

WebGPU

MacBook Pro (M5 Max, 128 GB)

128x128

Real BF30 parity

Radius 30 px

Partial

WebGPU

MacBook Pro (M5 Max, 128 GB)

256x256

Deterministic fixture

Test BF

Test

WebGPU

MacBook Pro (M5 Max, 128 GB)

512x512

Real interaction

Frozen phase reference

Refuted

WebGPU

MacBook Pro (M5 Max, 128 GB)

1024x1024

Real interaction

Incomplete frozen reference

Partial

CPU reference

Portable CI runner

128x128

—

—

Not retained

CPU reference

Portable CI runner

256x256

—

—

Not retained

CPU reference

Portable CI runner

512x512

Independent adjudication

Frozen fixture

Ref

CPU reference

Portable CI runner

1024x1024

—

—

Not retained

Native Swift/Metal has a package-owned 512×512 implementation. Other native Swift scan sizes remain unsupported rather than inferred from CUDA/MPS. Untimed CUDA and MPS sizes retain fixed-size parity coverage. The WebGPU 512×512 phase result is explicitly refuted; its timing cannot be promoted until parity is restored.

See the SSB performance history for size-specific historical experiments. Current numerical rows remain in this dashboard and the verified-results ledger.

Cross-module platform map#

A one-row-per-platform map hides the computer, configuration, and evidence state, and previously made a refuted WebGPU SSB result look accepted. It is no longer maintained. Use the filterable atomic coverage matrix, where every row begins with Platform and Computer and keeps module, bin, dtype, cache state, and next gate separate.

A warm resident kernel, prepared source, first-process application load, and saved-result reopen answer different questions. See Benchmark methodology and Verified benchmark results before comparing them.

The public scientific array is always

\[ I[R_r,R_c,k_r,k_c], \]

where \(\mathbf R=(R_r,R_c)\) is the probe/scan coordinate and \(\mathbf k=(k_r,k_c)\) is the detector coordinate. Runtime-specific layout, tiling, fusion, and dispatch are private optimizations; shape, sampling, precision, calibration, and provenance are shared contracts.

Where an implementer starts#

Goal

Read first

Then inspect

Acceptance evidence

Change scientific meaning or add an operation

Scientific contract

Scientific kernels

Operation-specific equation, provenance schema, and independent reference

Optimize CUDA

CUDA implementation

The domain’s cuda/ package, for example resident/cuda

Real NVIDIA profile plus exact/frozen parity

Optimize Python on Apple Silicon

Python MPS

The domain’s mps/ package, for example resident/mps

Physical Apple device profile plus exact/frozen parity

Build a native Apple client/library

Native Swift and Metal

Package.swift and native/swift/{Sources,Tests}

swift test, physical Metal timing, and cross-language fixtures

Optimize a browser client

WebGPU

Domain WebGPU TypeScript/WGSL resources

Real adapter, headed browser gate, and matching scientific output

Deploy CUDA behind a process boundary

QuantEM.GPU Remote

Deployment, protocol, and admission pages

Same array/provenance contract plus transport and capacity checks

Add or review a benchmark

Benchmark methodology

Continuous profiling, parity, and the optimization ledger

Date, revision, device, source plan, cache state, memory, wall boundary, and parity artifact

Dashboard maintenance rule#

Update a dashboard row only after its detailed evidence row is complete. The detail remains authoritative and must record measurement date, exact source revision, physical device/runtime, source shape and dtype, cache state, crop/bin/load plan, benchmark definition, peak memory or swap where available, and numerical or hash parity. Keep an older result when the newer experiment changes any of those conditions; label both instead of silently replacing one. Keep every row atomic: a different bin, dtype path, fixture, cache state, or statistic is another row. Documentation tests enforce that the landing page stays timing-free and that current overview values remain owned by this dashboard.

Accepted and rejected experiments remain in the optimization ledger, and the machine-readable evidence fingerprints are in performance/evidence_manifest.json. The platform/module cadence, runner ownership, and known harness gaps are machine-checked from benchmarks/profile_matrix.json; see Continuous profiling.