Kernel and benchmark dashboard#
This is the one-page technical overview of quantem.gpu: what the scientific
kernels compute, where each runtime implements them, how parity is proved, and
what the latest retained measurements actually mean.
Read the state before comparing the number
First-process source load, prepared-source load, warm resident compute, and saved-result reopen are different experiments. A binned or cropped source is never presented as native resolution. Check the complete provenance ledger before using a number in a design or release decision.
Dashboard review: 2026-08-22. Every current measured timing below shows the device, test date, and exact source revision. The overview omits opaque evidence IDs and does not replace the complete benchmark provenance ledger.
Coverage and next runs#
The filterable coverage registry now keeps every required configuration visible, including configurations that have never run. It separates complete measurements, partial evidence, pending work, refuted experiments, and fail-closed unsupported paths. Each open row names a stable runbook, its physical owner, and the exact artifact required for promotion.
State |
Gate count |
|---|---|
✓ Measured |
18 |
◐ Partial |
34 |
○ Pending |
73 |
! Blocked |
35 |
× Refuted |
4 |
Not supported |
27 |
↺ Superseded |
0 |
Platform |
Tracked gates |
|---|---|
CUDA |
17 |
Python MPS |
37 |
Native Swift/Metal |
56 |
WebGPU |
58 |
CPU reference |
23 |
Platform and computer coverage#
Each row identifies one reproducible hardware configuration. Counts describe tracked cells, including explicit unsupported contracts; a pending value remains a test to run. Load, admission, memory, and performance gates are multiplied across compatible computers because hardware changes the result. Platform-wide correctness or unsupported contracts are recorded once instead of creating misleading duplicate hardware rows.
Platform |
Computer |
Tracked cells |
Measured |
Partial |
Pending |
Blocked |
Refuted |
Unsupported |
|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
17 |
2 |
5 |
5 |
0 |
0 |
5 |
Python MPS |
MacBook Air (M2, 8 GB) |
9 |
0 |
0 |
7 |
2 |
0 |
0 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
18 |
2 |
3 |
6 |
0 |
2 |
5 |
Python MPS |
MacBook Pro (M5, 24 GB) |
10 |
0 |
1 |
8 |
0 |
1 |
0 |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
14 |
0 |
0 |
11 |
2 |
0 |
1 |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
27 |
8 |
5 |
7 |
0 |
0 |
7 |
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
15 |
4 |
3 |
7 |
0 |
0 |
1 |
WebGPU |
MacBook Air (M2, 8 GB) |
15 |
0 |
0 |
4 |
11 |
0 |
0 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
27 |
2 |
4 |
7 |
9 |
1 |
4 |
WebGPU |
MacBook Pro (M5, 24 GB) |
16 |
0 |
0 |
7 |
9 |
0 |
0 |
CPU reference |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
2 |
0 |
0 |
2 |
0 |
0 |
0 |
CPU reference |
MacBook Air (M2, 8 GB) |
2 |
0 |
0 |
0 |
2 |
0 |
0 |
CPU reference |
MacBook Pro (M5 Max, 128 GB) |
2 |
0 |
2 |
0 |
0 |
0 |
0 |
CPU reference |
MacBook Pro (M5, 24 GB) |
2 |
0 |
0 |
2 |
0 |
0 |
0 |
CPU reference |
Portable CI runner |
15 |
0 |
11 |
0 |
0 |
0 |
4 |
Agents and maintainers should begin with:
python scripts/benchmark_registry.py next --limit 10
python scripts/benchmark_registry.py command GATE_ID
The overview tables below retain current measurements, qualified probes, and clearly marked diagnostics. Read the State cell before the time. The coverage registry is the canonical place to find missing combinations and their reproduction entry points.
Speed and memory at a glance#
The rows below are deliberately not a leaderboard: results with different fixtures, cache states, scientific plans, or wall-clock boundaries are not ranked against one another. Each row keeps those conditions together. Measurement tables are keyed first by Platform, then by the reproducible Computer class; local host nicknames are never public benchmark identifiers.
Current qualified load measurements#
One row is one exact configuration on one reproducible computer. Platform and computer are the first two columns; scan geometry, detector geometry, bin, dtype, cache state, statistic, memory observations, device, date, and revision remain separate data. The table is generated from the benchmark registry so a superseded result cannot remain the dashboard headline.
Only explicitly designated current rows appear here. Historical, superseded, refuted, and unmeasured configurations remain in the complete tables below rather than being silently deleted.
Platform |
Computer |
State |
Selected scan |
Source detector |
Detector bin |
Output detector |
Source dtype |
Staging dtype |
Resident dtype |
Scientific gate |
Cache/process state |
Wall boundary |
Samples |
p50 |
p95 |
Maximum |
Logical resident |
Accelerator/driver peak |
Process/tree peak |
Process physical-footprint peak |
Swap delta |
Parity |
Device tested |
Date tested |
Revision |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
uint16 |
uint16 |
exact |
one same-process warmup; operating-system source pages uncontrolled; no eviction performed; fresh returned destination released after each trial |
public io.load return after MPS synchronization; full-volume hash and release excluded |
7 |
0.406624 s |
0.428164 s |
0.428164 s |
18.000 GiB |
18.442 GiB |
18.692 GiB |
n/a |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
uint8 |
uint16 |
exact |
fresh QH5 index root and destination; macOS F_NOCACHE on source hashing and indexed descriptors; immutable source already audited |
catalog, pipeline compilation, plan, complete exact private-resident volume, seven products, metadata, and provenance |
7 |
0.577793 s |
0.900979 s |
0.900979 s |
18.000 GiB |
18.571 GiB |
0.874 GiB |
n/a |
n/a |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
n/a |
uint16 |
n/a |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.225984 s |
0.231192 s |
0.231192 s |
4.500 GiB |
5.087 GiB |
0.657 GiB |
5.336 GiB |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
n/a |
uint16 |
n/a |
prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded |
6 |
0.200036 s |
0.204694 s |
0.204694 s |
1.125 GiB |
1.712 GiB |
1.039 GiB |
n/a |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
n/a |
uint16 |
exact |
original compressed source reread every visit; existing QH5 index and source-bound packing layout; exact DPC sums reused after first load; OS pages uncontrolled |
indexed source open through synchronous complete packed-resident return; catalog and independent full-count audit separate |
7 |
1.496660 s |
2.086119 s |
2.086119 s |
18.000 GiB |
n/a |
n/a |
2.531 GiB |
n/a |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-09-06 |
|
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
n/a |
uint16 |
n/a |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.678703 s |
0.697179 s |
0.697179 s |
4.500 GiB |
5.087 GiB |
0.662 GiB |
5.271 GiB |
0 B |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-08-22 |
|
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
n/a |
uint16 |
n/a |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.640942 s |
0.651931 s |
0.651931 s |
1.125 GiB |
1.712 GiB |
1.018 GiB |
2.319 GiB |
0 B |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-08-22 |
|
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
uint16 |
uint16 |
exact |
prepared immutable block indexes; explicitly warm source pages; fresh browser target per retained run; 699 pageouts; zero swap growth; no dropped runs |
navigation through scientifically usable exact resident output and diagnostic frame checksums; exhaustive full-volume hash and application E2E excluded |
7 |
1.358000 s |
1.594000 s |
1.594000 s |
18.000 GiB |
n/a |
6.500 GiB |
n/a |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false |
2026-08-22 |
|
CPU reference |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
uint16 |
uint16 |
exact |
prepared indexes; source pages unspecified; CPU exact-reference validation only |
public CPU reference load; full-volume hash excluded |
1 |
32.788039 s |
32.788039 s |
32.788039 s |
18.000 GiB |
n/a |
19.141 GiB |
n/a |
0 B |
Pass |
Apple M5 Max CPU; 128 GB unified memory |
2026-08-22 |
|
These measurements use a complete 512x512 scan and native
192x192 uint16 source detector, with scan bin 1 and no crop. The current
Python MPS, Native Swift/Metal, and CPU rows use fixture
real-512x512x192x192-u16-bslz4-27shard-master-fixture-c. The physical WebGPU
smoke uses the separately fingerprinted
full-native-webgpu-512x512x192x192-u16 fixture; it is not a cross-fixture
speed comparison. Every row retains its own fixture SHA-256 and source identity
in the complete registry. Detector bin is explicit. Exact binned uint16
output is source-identity-bound to its retained maximum-count audit; it is not
a general license to narrow arbitrary input.
For this immutable source, the complete maximum-count audit of 53 proves that
an 8x8 exact detector sum is at most 3,392, so bins through 8 fit in uint16.
Other sources must pass their own complete range audit or use a provably
sufficient wider integer dtype.
The package rows are prepared-index measurements. Python MPS source pages were
uncontrolled after one same-process warmup. Native Swift/Metal uses prepared
immutable indexes and controlled F_NOCACHE source descriptors. Neither is a
cold arbitrary-source or application end-to-end result. Historical CUDA,
WebGPU, earlier CPU, destination-reuse, and superseded measurements remain in
Verified benchmark results and the complete
coverage registry; they are not silently discarded
or mixed into this current table.
Logical resident bytes, accelerator/driver peak, process RSS high-water,
process physical-footprint peak, and swap delta are distinct observations and
are not additive. A missing physical-footprint value is displayed as n/a,
not inferred from RSS. Process RSS does not include every direct Metal
allocation, so it cannot replace accelerator or physical-footprint telemetry.
For example, the current full-native Python MPS row reports 18.442 GiB of
driver allocation for an 18.00 GiB logical resident tensor; that is a measured
after-load boundary, not a universal peak estimate.
On the 24 GB MacBook Pro, the full-native bin-1 smoke
passed every exact volume and product hash but produced approximately
723.31 MiB additional swap use and 758,448,128 B of swapouts. The safety gate
therefore stopped after one trial; it is partial evidence, not a repeatable
performance distribution.
Dtype support and peak memory#
“Source,” “staging,” “accumulation,” and “resident” dtype describe different
stages. A uint8 row is scientifically exact only when the source is already
uint8 or a source-identity-bound complete audit proves
maximum <= 255 and pixelsAbove255 == 0. Otherwise a saturating uint8
output, such as the WebGPU clip8 browse decode, clips values above 255 and is a
browse representation, not raw-count evidence.
The registry below separates those cases. Each row fixes one platform, computer class, scan, detector, bin, crop, source dtype, staging dtype, resident dtype, and scientific gate. Browse-only can be benchmarked, but it cannot satisfy an exact scientific gate.
Each row is one atomic source, staging, and resident dtype contract. Exact rows preserve scientific counts; browse-only rows are explicit saturating representations and cannot satisfy an exact gate.
Platform |
Computer |
State |
Selected scan |
Source detector |
Scan bin |
Detector bin |
Output detector |
Crop |
Source dtype |
Staging dtype |
Resident dtype |
Scientific gate |
Precision contract |
Implementation basis |
Next gate or reason |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint8 |
uint8 |
uint8 |
exact |
native integer counts |
src/quantem/gpu/io/hdf5/cuda/kernels/bslz4.cu shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=cuda) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained. |
Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public CUDA io.load and retain its complete-volume hash, source/resident provenance and peak memory. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint16 |
exact |
native integer counts |
The CUDA decoder has a dedicated uint16 bitshuffle path and existing unit parity; the current registry has no uncontended full-volume native uint16 distribution. |
Run the public CUDA load entry point on an uncontended device and retain a complete-volume uint16 hash plus source/staging/resident provenance and peak allocation. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint32 |
uint32 |
uint32 |
exact |
native integer counts |
src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form. |
CUDA io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint8 |
exact |
source-identity-bound complete value-range audit |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output. |
CUDA io.load has no uint8 output dtype. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint8 |
browse-only |
explicit saturation to 255 |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output. |
CUDA io.load has no saturating uint8 browse output. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint32 |
exact |
lossless integer widening |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening. |
CUDA io.load has no uint16-to-uint32 widening. |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint16 |
exact |
source-identity-bound uint8 staging with exact uint16 reconstruction |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging. |
Audited uint8 staging with uint16 resident reconstruction is not implemented for CUDA. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint8 |
uint8 |
uint8 |
exact |
native integer counts |
src/quantem/gpu/io/hdf5/mps/kernels/bslz4.msl shuf_8_batched unshuffles one-byte bitshuffle/LZ4 blocks, including the partial final block; tests/hardware/test_bitshuffle_uint8.py (QEM_TEST_BACKEND=mps) matches h5py for full, partial and multi-block frames. No real uint8 acquisition is retained. |
Load a real full-volume native uint8 bitshuffle/LZ4 acquisition through public Python MPS io.load and retain its complete-volume hash, source/resident provenance and peak memory. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint16 |
exact |
native integer counts |
The MPS uint16 decoder and retained full-volume canonical hash preserve native counts at detector bin 1. |
n/a |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint32 |
uint32 |
uint32 |
exact |
native integer counts |
src/quantem/gpu/io/encoded.py holds a uint32 source as exact uint16 codes when every count fits 16 bits and rejects it otherwise; the encoded resident has no uint32 form. |
Python MPS io.load has no uint32 resident: uint32 counts that fit 16 bits are held exactly as uint16, larger ones are refused. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint8 |
exact |
source-identity-bound complete value-range audit |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint8 output. |
Python MPS io.load has no uint8 output dtype. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint8 |
browse-only |
explicit saturation to 255 |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no saturating uint8 browse output. |
Python MPS io.load has no saturating uint8 browse output. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint32 |
exact |
lossless integer widening |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no uint32 widening. |
Python MPS io.load has no uint16-to-uint32 widening. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint16 |
exact |
source-identity-bound uint8 staging with exact uint16 reconstruction |
src/quantem/gpu/io/load.py accepts dtype only as native (or scaled_uint16 for float32 sources): io.load keeps stored counts and has no audited uint8 staging. |
Audited uint8 staging with uint16 resident reconstruction is not implemented for Python MPS. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint8 |
uint8 |
uint8 |
exact |
native integer counts |
Native4DSTEMIO can describe one-byte indexed sources, but Metal4DSTEMIndexedLoader.swift requires a uint16 source and uint16 staging at the public integrated boundary. |
The integrated indexed Swift/Metal loader does not admit a native uint8 source. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint16 |
exact |
native integer counts |
Metal4DSTEMIndexedLoader.swift supports the uint16 fallback stage, while retained optimized measurements used separately registered audited uint8 staging. |
Run a source whose complete audit does not authorize uint8 staging and retain exact uint16 staging-to-resident parity and allocation evidence. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint32 |
uint32 |
uint32 |
exact |
native integer counts |
NativeHDF5Bridge.swift describes one- and two-byte indexed sources, and Metal4DSTEMIndexedLoader.swift requires uint16 input. |
The native indexed Swift/Metal source contract does not admit uint32 HDF5 detector values. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint8 |
exact |
source-identity-bound complete value-range audit |
Metal4DSTEMExactBinner.provenance accepts only uint16 or uint32 output, although it can use an audited uint8 staging buffer. |
The exact native Swift/Metal load contract cannot publish a uint8 resident scientific volume. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint8 |
browse-only |
explicit saturation to 255 |
The indexed exact loader exposes uint16 or uint32 scientific output and has no separate saturating browse-resident API. |
Explicit saturating uint8 resident output is not implemented by the public native Swift/Metal load boundary. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint32 |
exact |
lossless integer widening |
Metal4DSTEMExactBinner has tested uint16-to-uint32 kernels and provenance, but Metal4DSTEMIndexedBinnedLoad fixes integrated resident output to uint16. |
Expose uint32 output through the public indexed loader and cache contract, then retain complete load, metadata, and physical parity. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint16 |
exact |
source-identity-bound uint8 staging with exact uint16 reconstruction |
Native4DSTEMValueRangeAudit binds source identity and maximum counts; the retained controlled full-native run used audited uint8 staging and reproduced exact uint16 volume and product hashes. |
n/a |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint8 |
uint8 |
uint8 |
exact |
native integer counts |
h5reader.ts, bslz4.ts, and local-h5.ts recognize matching native uint8 source, decode, and resident modes; no retained full-volume physical uint8 fixture proves the complete path. |
Run an exact physical hardware-browser full-volume uint8 source through local HDF5 decode and retain output hash, allocation, and source/decode/resident provenance. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint16 |
exact |
native integer counts |
The physical WebGPU path retained exact native uint16 resident checks and seven runs, but device allocation and timed complete-volume hashing remain incomplete. |
Add WebGPU device-allocation telemetry and retain an exact complete-volume hash inside the timed boundary before promoting the integrated load gate. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint32 |
uint32 |
uint32 |
exact |
native integer counts |
h5reader.ts, bslz4.ts, and local-h5.ts implement matching native uint32 source/decode/resident modes; physical full-volume proof is absent. |
Run a native uint32 HDF5 source on physical hardware WebGPU and retain the complete resident hash, browser/device memory, and source/decode/resident provenance. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint8 |
exact |
source-identity-bound complete value-range audit |
bslz4.ts has an audit-dependent low8 kernel, but local-h5.ts does not bind a complete audit identity to the returned uint8 resident state. |
Replace the experimental low8 global flag with a typed source-identity-bound audit in the local-HDF5 public result, then retain exact full-volume parity. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint8 |
browse-only |
explicit saturation to 255 |
bslz4.ts decodes all uint16 planes and saturates to 255; source tests distinguish this from the experimental low8 audit path. |
Run the fused full-plane clip8 path on physical hardware with values above 255 and retain output hash, browser/device memory, and browse-only provenance. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint32 |
exact |
lossless integer widening |
local-h5.ts requires decode dtype uint32 to match a uint32 HDF5 source and rejects a uint16 source request. |
The WebGPU local-HDF5 public path does not widen uint16 source values into uint32 resident storage. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint16 |
exact |
source-identity-bound uint8 staging with exact uint16 reconstruction |
local-h5.ts keeps decode and resident integer mode matched for native paths and rejects mismatched uint32 requests; it exposes no uint8-stage-to-uint16 reconstruction contract. |
Audited uint8 staging followed by exact uint16 resident reconstruction is not implemented in the public WebGPU local-HDF5 path. |
CPU reference |
Portable CI runner |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint8 |
uint8 |
uint8 |
exact |
native integer counts |
src/quantem/gpu/io/hdf5/cpu.py returns the HDF5 native dtype unchanged at detector bin 1; a complete public-API uint8 fixture artifact is not retained. |
Add a repository fixture that proves public io.load preserves every native uint8 count, shape, order, and provenance field at detector bin 1. |
CPU reference |
Portable CI runner |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint16 |
exact |
native integer counts |
The CPU reference returns native uint16 counts at detector bin 1 and has retained physical reference probes; the portable registry gate is not a complete public-API artifact. |
Promote only after the portable public API fixture retains complete-volume uint16 hash, shape, order, mask, and metadata parity. |
CPU reference |
Portable CI runner |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint32 |
uint32 |
uint32 |
exact |
native integer counts |
The CPU reference preserves the HDF5 native dtype at detector bin 1; public load defaults may auto narrow uint32, so the exact uint32 request must be explicit and tested. |
Add a public-API native uint32 fixture that disables advisory auto narrowing and proves complete count, dtype, shape, and metadata parity. |
CPU reference |
Portable CI runner |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint8 |
exact |
source-identity-bound complete value-range audit |
src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint8 output. |
CPU reference io.load has no uint8 output dtype. |
CPU reference |
Portable CI runner |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint8 |
browse-only |
explicit saturation to 255 |
src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no saturating uint8 browse output. |
CPU reference io.load has no saturating uint8 browse output. |
CPU reference |
Portable CI runner |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint16 |
uint32 |
exact |
lossless integer widening |
src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no uint32 widening. |
CPU reference io.load has no uint16-to-uint32 widening. |
CPU reference |
Portable CI runner |
Not supported |
512 × 512 |
192 × 192 |
1 |
1 |
192 × 192 |
none |
uint16 |
uint8 |
uint16 |
exact |
source-identity-bound uint8 staging with exact uint16 reconstruction |
src/quantem/gpu/io/load.py refuses every dtype request on the CPU reference, which returns stored counts and has no audited uint8 staging. |
Audited uint8 staging with uint16 resident reconstruction is not implemented for CPU reference. |
This capability matrix does not infer peak memory. The current load table is authoritative only for its matching shape, dtype, device, cache state, and wall boundary. In particular, browser RSS may be smaller than a WebGPU resident payload because it does not capture every device allocation; that is an incomplete peak, not evidence that the payload disappeared.
The current C matrix measures native uint16 input. The full-native logical
payload is exactly 19,327,352,832 bytes (18.00 GiB). Python io.load on CUDA
and MPS keeps that complete native detector and holds it ANS encoded, so its
resident bytes depend on the acquisition, not on a detector bin. Detector bins
are explicit options of the native Swift/Metal and WebGPU loaders and exact
views of the encoded resident in the CUDA browse service. The native
Swift/Metal primitive has admitted one full 18.00 GiB
resident load on the 24 GB Apple M5, but repeated uncontended memory and paging
evidence remains incomplete. A consumer may apply a conservative working-set
limit; that application policy is separate from primitive capability and may
not silently select detector bin 2.
The current production WebGPU detector-bin-2/4/8 path accumulates into and
stores float32. Those historical timings remain useful implementation
history, but they do not satisfy the exact-integer resident contract. The
platform/computer registry therefore marks exact WebGPU bins 2/4/8 blocked
until integer accumulation and residency pass full-volume parity.
Seven-run physical exact full-volume distribution evidence remains in the
retained-measurement table; unavailable WebGPU device-allocation telemetry is
still shown as n/a, never inferred from browser RSS.
Public Python io.load keeps stored counts: omit dtype or pass
dtype="native". Its one conversion is an explicit dtype="scaled_uint16" for
float32 sources, which stores calibrated integer codes, stays ANS encoded, and
reports the measured conversion error. The CUDA, Python MPS and CPU reference
rows above for uint8 outputs, uint32 widening and audited uint8 staging are
therefore Not supported: those were load-time casts that io.load no
longer offers.
Minimum-device memory gates#
The public release floors are 6 GiB of dedicated VRAM for CUDA and 8 GB
of total laptop RAM for WebGPU. The WebGPU number is the entire machine
budget shared by the operating system, browser, JavaScript heap, staging
buffers, and GPU—not memory available exclusively to one GPUBuffer.
✓ is awarded only after the complete load-and-product pipeline runs on a physical device at or below the stated floor with retained peak memory, pressure/swap, output parity, and responsiveness evidence. A calculated payload can prove No, but it can only establish a Pending candidate. An allocator cap on a larger GPU is a useful pre-check, not physical-device signoff.
Platform |
Computer |
Minimum device |
Selected scan |
Scan plan |
Source detector |
Detector bin |
Output detector |
Resident dtype |
Resident payload |
Gate |
Reason |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Physical 6 GiB CUDA computer pending |
6 GiB VRAM |
|
Full |
|
1 |
|
|
1.98 to 3.14 GiB |
Pending |
Measured encoded payload of three real acquisitions; complete physical 6 GiB peak not retained |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB unified RAM |
|
Full |
|
1 |
|
|
Data-dependent |
Pending |
Encoded payload and physical 8 GB pressure/parity run not retained |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB unified RAM |
|
Full |
|
1 |
|
|
18.00 GiB |
Blocked |
Resident payload exceeds physical memory |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB unified RAM |
|
Full |
|
2 |
|
|
4.50 GiB |
Pending |
Current clean-revision pressure/parity run not retained |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB unified RAM |
|
Full |
|
4 |
|
|
1.125 GiB |
Historical |
Earlier physical run passed; current clean-revision repeat remains pending |
Apple M2 MacBook Air ( |
2026-08-18 |
|
MacBook Air (M2, 8 GB) |
8 GB unified RAM |
|
Full |
|
8 |
|
— |
— |
Not supported |
Current native load-plan contract supports bins 1, 2, and 4 |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB total RAM |
|
Full |
|
1 |
|
|
18.00 GiB |
Blocked |
Resident payload exceeds total machine RAM |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB total RAM |
|
Full |
|
2 |
|
|
4.50 GiB |
Blocked |
Production detector binning stores |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB total RAM |
|
Full |
|
4 |
|
|
1.125 GiB |
Blocked |
Production detector binning stores |
— |
— |
|
MacBook Air (M2, 8 GB) |
8 GB total RAM |
|
Full |
|
8 |
|
|
0.28125 GiB |
Blocked |
Production detector binning stores |
— |
— |
Python io.load on CUDA and MPS has no detector bin: it keeps the complete
native detector ANS encoded, so its payload depends on the counts. Three real
512x512x192x192 uint16 acquisitions encode to 2,127,983,353,
2,279,376,159 and 3,374,307,987 bytes (0.11 to 0.18 of the 18.00 GiB of
counts) on the CUDA loader (NVIDIA RTX PRO 6000 Blackwell, 2026-10-05). The
exact integer WebGPU bin-2/4/8 payloads can fit within the nominal 8 GB
total-RAM floor, but payload arithmetic alone is not acceptance. Production WebGPU still
needs integer accumulation/residency plus a physical browser run that captures
browser, adapter, staging, operating-system, pressure, swap, and scientific
parity. Likewise, the current Blackwell timings do not prove a physical 6 GiB
CUDA floor.
What a 4 or 6 GiB budget can hold#
This capacity chart fixes the full scan at 512x512 and the native detector at
192x192. “Payload” excludes decoder scratch, staging buffers, allocator
reserve, and other GPU users unless the row reports a measured process peak.
Detector bin |
Resident dtype |
Resident payload |
4 GiB fit |
6 GiB fit |
Evidence state |
|---|---|---|---|---|---|
1 |
|
18.00 GiB |
No |
No |
Calculated dense payload (native Swift/Metal and WebGPU bin 1) |
1 |
|
1.98 to 3.14 GiB |
Candidate |
Candidate |
Measured CUDA |
1 |
|
9.00 GiB |
No |
No |
Complete-audit lossless path only (WebGPU low8) |
4 |
|
1.125 GiB |
Candidate |
Candidate |
Physical 8 GB M2 Air evidence (native Swift/Metal); 4/6 GiB signoff Pending |
The encoded CUDA/MPS load needs no memory plan of its own: io.load reads and
encodes the acquisition in bounded scan blocks, and read(...) and the
detector kernels decode bounded blocks, so the working set is the encoded
payload plus bounded decode scratch.
Each 512x512 float32 product map is only 1 MiB, and one 192x192
float32 mean diffraction pattern is 144 KiB. The source working set—not the
final BF/ADF/DF/DPC image—is the capacity problem.
For a full scan with output detector bin \(b\) and \(w\) resident bytes per value, the payload alone is
Peak memory is larger: the benchmark must also report live compressed bytes, decode and reduction scratch, staging/upload buffers, allocator reserve, products, concurrent GPU users, and—on unified memory—process pressure and swap. A calculated payload is never relabeled as a measured peak.
Small-GPU support today
The CUDA/MPS screening.prepare path loads the complete acquisition once into
encoded storage and emits mean DP, total, BF, ABF, ADF, DF, CoM, rotation, and
iDPC without cropping the scan. Physical 4 and 6 GiB product-pipeline signoff
is Pending. See the screening API before
choosing a plan.
Platform-first module dashboard#
Every table keeps the scientific module as its section and puts the execution platform in the first column. Empty cells are forbidden:
✓ — verified with retained real-data or physical-device parity evidence.
Test — deterministic source/test coverage without an equivalent retained physical or native-data run.
Pending — implementation exists, but that exact size, bin, or timing evidence has not been retained yet.
Ref — CPU correctness adjudication, never a production fallback.
— — unsupported or not a target.
I/O and first usable product — quantem.gpu.io#
Exact configuration gaps#
The measured table above owns retained full-scan timings. The filterable coverage registry owns every missing combination as one Platform → Computer → detector-bin → cache-state row; this page does not maintain a second bundled list. Real-space crop remains an explicit selector and is never an implicit performance policy.
Selective scan rectangles#
The physical WebGPU rows below are prepared whole-shard-selective rectangle loads. A retained frame-span manifest lets the loader omit nonintersecting shards and decode/upload only selected scan rows. It does not issue byte-range reads inside an intersecting shard.
Platform |
Computer |
Selected scan |
Rectangle |
Source detector |
Source dtype |
Shards read |
Storage bytes read |
Samples |
Loader p50 |
Loader p95 |
Loader maximum |
Logical resident |
Browser-tree RSS peak |
Observed swap delta |
Parity |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
|
|
|
4 of 27 |
488,224,242 B |
5 |
0.147 s |
0.1544 s |
0.156 s |
301,989,888 B |
1,724,317,696 B |
0 B |
✓ |
Chrome 151, Apple M5 Max Metal |
2026-08-22 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
|
|
|
14 of 27 |
1,705,556,941 B |
5 |
0.381 s |
0.3924 s |
0.394 s |
4,831,838,208 B |
3,002,875,904 B |
0 B |
✓ |
Chrome 151, Apple M5 Max Metal |
2026-08-22 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
|
|
|
20 of 27 |
2,432,636,897 B |
5 |
0.574 s |
0.582 s |
0.584 s |
10,871,635,968 B |
3,896,934,400 B |
0 B |
✓ |
Chrome 151, Apple M5 Max Metal |
2026-08-22 |
All rows use prepared block indexes and a prepared frame-span manifest. The
operating-system source-page state was uncontrolled/unspecified and no eviction
was performed. Frame-manifest read/encoding, DevTools
injection, checksum harness, products, and application E2E are excluded. A
negative control without the frame-span manifest read all 27 shards (3.17 GB),
so it is explicitly not selective evidence. CUDA and Python MPS load the
complete acquisition and return a rectangle with
read(scan_region=..., detector_region=...) from the encoded resident; native
Swift/Metal bulk selection and physical cross-backend timing remain pending.
The parity check mark means all five runs for that rectangle passed independent
CPU uint16 checksum probes at the first, middle, and last retained raw frames,
plus output shape, dtype, row-major order, and selection metadata. Full-tensor
readback parity was not performed. The WebGPU fixture view records source
identity 1be810b9...; fixture C records c9c0d968..., and the CUDA master
record is 4802ec16.... These rows are not a cross-lane fixture-controlled
comparison.
Screening and prepared-product caches — quantem.gpu.screening#
Platform |
Computer |
Support |
Operation |
Source plan |
Statistic |
Time |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
✓ |
Exact streamed screening |
Full |
p50 of 6 |
1.205 s |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, GPU 0 |
2026-08-22 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
✓ |
Exact screening build |
Full |
Single run |
6.711 s |
Apple M5 Max ( |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
✓ |
Validated screening-v3 reopen |
Prepared derived products |
p50 |
20.803 ms |
Apple M5 Max ( |
2026-08-19 |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
✓ |
Validated exact-summary reopen |
Prepared derived products |
p50 |
0.029 s |
Apple M2 MacBook Air ( |
2026-08-19 |
WebGPU |
N/A (not implemented) |
— |
Prepared screening workflow |
— |
— |
— |
— |
— |
CPU reference |
Portable CI runner |
Ref |
Independent adjudication |
Reference fixtures |
— |
Pending |
— |
— |
The promoted MPS build derives its detector mean and masks from the full scan. When the provisional and final masks differ, BF/DF are recomputed exactly. The first-chunk-mask candidate failed parity and is retained only as a rejected experiment; its timing is not repeated on a current page. Saved-product reopen is likewise never presented as source load.
The current CUDA pinned-slot candidate reduced the like-for-like package p50
from 1.356516 s to 1.204713 s; candidate p95/max were
1.325760/1.329731 s. All six public arrays were byte exact in every trial.
This is warm/source-pages-unspecified exact screening, not cold HDF5 and not a
prepared-result reopen. Detailed pinned-registration and memory measurements
remain in the provenance ledger.
Virtual images — quantem.gpu.detector#
Complete platform, computer, and detector-bin matrix#
The registry keeps one row per exact product configuration instead of hiding an unrun combination. Use the Platform, Computer, State, and Detector bin filters to isolate the implementation or physical computer you care about. A pending, blocked, or unsupported row is intentionally visible; only a measured row publishes a timing distribution.
The 8 GB MacBook Air detector-bin-2 row is the locked downstream acceptance configuration. Detector bins 4 and 8 remain separate library-capability cells, not silent application fallbacks. Detector bin 1 is explicitly blocked where its 18 GiB exact resident input cannot be admitted.
Each row fixes one platform, reproducible computer, detector bin, and exact product-suite boundary. Partial diagnostics retain their sample and memory context, but only fully measured gates display p50, p95, and maximum timing.
Platform |
Computer |
State |
Selected scan |
Source detector |
Detector bin |
Output detector |
Resident dtype |
Cache/process state |
Wall boundary |
Samples |
p50 |
p95 |
Maximum |
Logical resident |
Accelerator/driver peak |
Process/tree peak |
Swap delta |
Parity |
Device tested |
Date tested |
Revision |
Next gate or reason |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
warm operating-system source pages; unspecified; empty screening-result cache |
quantem.gpu.screening.prepare package wall through exact complete six-array screening result in balanced A-B-B-A-B-A-A-B comparison |
6 |
n/a |
n/a |
n/a |
n/a |
2.084 GiB |
1.958 GiB |
0 B |
Pass |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition; GPU0; driver 580.159.03 |
2026-08-22 |
|
Run the current public ScreeningResult on an uncontended CUDA device with benchmark_screening.py –require-exact-full-suite; retain all nine hashes, uint64 total/ABF/ADF dtype and shape, p50/p95/maximum, and complete memory telemetry. |
Python MPS |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
not admissible |
blocked exact-residency contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before decoder, allocator, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming product path that never requires an 18 GiB resident input before attempting this physical row. |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
exact resident detector-bin-1 MPS input; source pages uncontrolled or warm; no prepared product result; five pageouts; zero swapout growth |
single first-execution diagnostic only; synchronized detector-product publication was 0.499909 seconds; no p50, p95, or maximum distribution |
1 |
n/a |
n/a |
n/a |
18.000 GiB |
27.510 GiB |
20.436 GiB |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
Repeat the exact product suite without pageout growth for synchronized p50, p95, and maximum timing plus complete memory telemetry. |
Python MPS |
MacBook Pro (M5, 24 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
resident exact full-native MPS input |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run only after an uncontended pressure preflight; use a single exact full-native smoke with a hard stop on swap/pageout growth before considering repeated product timing. |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
not admissible |
blocked exact-residency contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
The exact 18 GiB private-resident input exceeds the computer’s 8 GB unified memory before decoder, Metal, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming or mapped product path that never requires an 18 GiB private-resident input before attempting this physical row. |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
○ Pending |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
resident exact detector-bin-2 Metal input |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run the complete product suite from the explicit exact detector-bin-2 resident volume with full parity, synchronized timing, pressure, swap, and peak memory. |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
○ Pending |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
resident exact detector-bin-4 Metal input |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run the complete exact detector-bin-4 product suite on the current clean revision with full hashes, synchronized timing, pressure, compressed memory, swap, and peak memory; detector bin 2 remains the locked application acceptance policy. |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
Not supported |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
not applicable |
unsupported load-plan contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not applicable |
n/a |
n/a |
n/a |
Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed. |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
prepared immutable QH5 indexes; source pages unspecified |
complete exact resident volume, products, metadata, calibration, and provenance; full private readback/hash excluded |
6 |
0.313870 s |
0.318865 s |
0.318865 s |
18.000 GiB |
18.571 GiB |
0.639 GiB |
n/a |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
n/a |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.225984 s |
0.231192 s |
0.231192 s |
4.500 GiB |
5.087 GiB |
0.657 GiB |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
n/a |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
prepared immutable indexes; F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven products; source discovery, audit creation, first allocation, pipeline compilation, application presentation, and post-boundary hashes excluded |
6 |
0.200036 s |
0.204694 s |
0.204694 s |
1.125 GiB |
1.712 GiB |
1.039 GiB |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; 128 GB unified memory |
2026-08-22 |
|
n/a |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Not supported |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
not applicable |
unsupported load-plan contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not applicable |
n/a |
n/a |
n/a |
Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed. |
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
prepared QH5 indexes; F_NOCACHE source descriptors; operating-system source-page state unspecified |
package exact indexed load plus exact full-scan products; app UI, index creation, and full-volume hash excluded |
1 |
n/a |
n/a |
n/a |
18.000 GiB |
18.571 GiB |
0.624 GiB |
n/a |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-08-22 |
|
Repeat uncontended full-native publication with complete product hashes, p50/p95/max, driver allocation, RSS, pressure, and swap. |
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.678703 s |
0.697179 s |
0.697179 s |
4.500 GiB |
5.087 GiB |
0.662 GiB |
0 B |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-08-22 |
|
n/a |
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
prepared immutable indexes; controlled F_NOCACHE source descriptors; compiled pipeline; exact private-resident destination reused for retained trials |
exact indexed load, detector sum, and seven native-resolution products; source discovery, audit creation, pipeline compilation, application presentation, and separately timed boundary hashes excluded |
6 |
0.640942 s |
0.651931 s |
0.651931 s |
1.125 GiB |
1.712 GiB |
1.018 GiB |
0 B |
Pass |
Apple M5 10-core integrated GPU; 24 GB unified memory |
2026-08-22 |
|
n/a |
Native Swift/Metal |
MacBook Pro (M5, 24 GB) |
Not supported |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
not applicable |
unsupported load-plan contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not applicable |
n/a |
n/a |
n/a |
Metal4DSTEMLoadPlan currently supports detector bins 1, 2, and 4 only; detector bin 8 must fail closed. |
WebGPU |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
not admissible |
blocked exact-residency contract |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
The exact 18 GiB resident input exceeds the computer’s 8 GB total memory before browser, adapter, decoder, products, and operating-system headroom. Next: Implement and prove an exact bounded-streaming browser product path that never requires an 18 GiB resident input before attempting this physical row. |
WebGPU |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
resident exact detector-bin-2 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory. |
WebGPU |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
resident exact detector-bin-4 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory. |
WebGPU |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
resident exact detector-bin-8 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
✓ Measured |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
prepared QH5 indexes; warm source pages; exact full-native uint16 resident input; fresh browser target per retained run |
complete internal GPU product suite from exact resident input through awaited readbacks; parity hashing, sorting, and encoding excluded |
7 |
0.482000 s |
0.483600 s |
0.483600 s |
18.000 GiB |
n/a |
6.492 GiB |
0 B |
Pass |
Apple M5 Max 40-core integrated GPU; Chrome 151; Apple Metal-3 hardware adapter; software=false |
2026-08-22 |
|
n/a |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
! Blocked |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
resident exact detector-bin-2 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
! Blocked |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
resident exact detector-bin-4 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory. |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
! Blocked |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
resident exact detector-bin-8 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory. |
WebGPU |
MacBook Pro (M5, 24 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
resident exact full-native hardware-WebGPU input |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run a guarded physical hardware-browser admission smoke, then retain the complete native product suite with full parity, timing, browser/device memory, pressure, and swap if admission remains safe. |
WebGPU |
MacBook Pro (M5, 24 GB) |
! Blocked |
512 × 512 |
192 × 192 |
2 |
96 × 96 |
uint16 |
resident exact detector-bin-2 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-2 accumulation and residency, then run the complete hardware product suite with full parity, timing, and browser/device memory. |
WebGPU |
MacBook Pro (M5, 24 GB) |
! Blocked |
512 × 512 |
192 × 192 |
4 |
48 × 48 |
uint16 |
resident exact detector-bin-4 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-4 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory. |
WebGPU |
MacBook Pro (M5, 24 GB) |
! Blocked |
512 × 512 |
192 × 192 |
8 |
24 × 24 |
uint16 |
resident exact detector-bin-8 hardware-WebGPU input required |
complete product suite publication |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
Production detector binning currently produces a float32 resident volume and cannot satisfy the exact-integer input contract. Next: Integrate exact-integer detector-bin-8 accumulation and residency, then run the complete hardware product suite with full parity, timing, pressure, swap, and browser/device memory. |
CPU reference |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
resident exact full-native CPU input |
complete reference product suite |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run the independent full-native CPU reference product suite only when host CPU, storage, and memory ownership are uncontended; retain all hashes, timing, and peak RSS. |
CPU reference |
MacBook Air (M2, 8 GB) |
! Blocked |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
not admissible |
blocked reference product suite |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Not run |
n/a |
n/a |
n/a |
The exact 18 GiB resident input exceeds the computer’s 8 GB unified memory before products and operating-system headroom. Next: Implement and prove an exact bounded-streaming CPU reference product path that never requires an 18 GiB resident input before attempting this physical row. |
CPU reference |
MacBook Pro (M5 Max, 128 GB) |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
exact full-native CPU-resident input after a warm operating-system page-cache source load |
single independent detector-product traversal was 31.078040 seconds; source load excluded; no p50, p95, or maximum distribution |
1 |
n/a |
n/a |
n/a |
18.000 GiB |
n/a |
36.450 GiB |
n/a |
Pass |
Apple M5 Max CPU; 128 GB unified memory |
2026-08-19 |
|
Add ABF, retain all product hashes, and repeat the exact physical CPU reference traversal for p50/p95/maximum and complete process-memory telemetry. |
CPU reference |
MacBook Pro (M5, 24 GB) |
○ Pending |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
resident exact full-native CPU input |
complete reference product suite |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Run only after an uncontended memory preflight; retain one exact full-native CPU reference smoke, complete product hashes, pressure, swap, and peak RSS before repetition. |
CPU reference |
Portable CI runner |
◐ Partial |
512 × 512 |
192 × 192 |
1 |
192 × 192 |
uint16 |
resident reference input |
complete reference product suite |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
n/a |
Pending |
n/a |
n/a |
n/a |
Keep exact reference tests current; no CPU speed claim is required. |
Retained operation-level timings#
Platform |
Computer |
Operation |
Scan grid |
Detector |
Detector bin |
Input state |
Statistic |
Time |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Mean diffraction |
|
|
1 |
Warm resident |
p50 |
18.392 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
BF exact sum |
|
|
1 |
Warm resident |
p50 |
3.768 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
ADF exact sum |
|
|
1 |
Warm resident |
p50 |
5.586 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
DF exact sum |
|
|
1 |
Warm resident |
p50 |
3.747 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Mean diffraction |
|
|
4 |
Warm resident |
p50 |
74.805 ms |
Apple M5 Max ( |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
BF exact sum |
|
|
4 |
Warm resident |
p50 |
2.502 ms |
Apple M5 Max ( |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
ADF exact sum |
|
|
4 |
Warm resident |
p50 |
4.404 ms |
Apple M5 Max ( |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
DF exact sum |
|
|
4 |
Warm resident |
p50 |
2.642 ms |
Apple M5 Max ( |
2026-08-19 |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Fused BF, ABF, ADF, total, and row/column moments |
|
|
1 |
Controlled source load |
p50 |
119.040 ms |
Apple M5 Max ( |
2026-08-22 |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
BF, ABF, ADF, total, and row/column moments |
|
|
4 |
Prepared resident-cache fallback |
Single run |
103.0 ms |
Apple M2 MacBook Air ( |
2026-08-19 |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
Exact moment widening for resident summary |
|
|
4 |
Same validated fused source pass; no resident traversal |
Single run |
0.569 ms |
Apple M2 MacBook Air ( |
2026-08-20 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Mean diffraction |
|
|
1 |
Warm resident |
p50 |
50.9 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
BF exact sum |
|
|
1 |
Warm resident |
p50 |
5.5 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
ADF exact sum |
|
|
1 |
Warm resident |
p50 |
15.0 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
DF exact sum |
|
|
1 |
Warm resident |
p50 |
43.4 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
CPU reference |
MacBook Pro (M5 Max, 128 GB) |
Virtual-image adjudication |
|
|
1 |
One independent traversal |
Single run |
31.08 s |
Apple M5 Max CPU |
2026-08-19 |
The current integer and mean-DP rows pass their independent CPU reference. CUDA
uses native detector resolution on fixture D; MPS uses explicit detector bin 4
on fixture C. WebGPU uses native detector resolution on D. The current native
bin-1 source pass produces three virtual images, detector sums, and exact
uint32 total and detector moments in one fused stage. That stage overlaps
source IO and is not additive with the package-wall measurement.
Revision 70bc366 proves when bin-4 accumulators fit, then widens the three
moment maps to uint64 in one small dispatch. The 0.569 ms row is only that
incremental widening; it is not the full source pass. The cache-only fallback
row remains because it has a different input state. Neither historical bin-4
row is a compressed-source load time.
Detector moments and phase contrast — quantem.gpu.dpc#
Platform |
Computer |
Operation |
Scan grid |
Detector |
Detector bin |
Input state |
Statistic |
Time |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
CoM row and column |
|
|
1 |
Warm resident |
p50 |
13.002 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Fixed-orientation iDPC |
|
|
1 |
CPU small-field integration |
p50 |
21.272 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
CoM row and column |
|
|
4 |
Warm resident |
p50 |
4.637 ms |
Apple M5 Max ( |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Fixed-orientation iDPC |
|
|
4 |
CPU small-field integration |
p50 |
12.678 ms |
Apple M5 Max ( |
2026-08-19 |
Native Swift/Metal |
MacBook Air (M2, 8 GB) |
CoM, DPC, iDPC, and display statistics |
|
|
4 |
Prepared exact |
Single run |
11.389 ms |
Apple M2 MacBook Air ( |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
DPC row |
|
|
1 |
Warm cached CoM |
p50 |
0.9 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
DPC column |
|
|
1 |
Warm cached CoM |
p50 |
0.7 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
iDPC |
|
|
1 |
Explicit 0-degree rotation |
p50 |
1.4 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
CPU reference |
MacBook Pro (M5 Max, 128 GB) |
Rotation and iDPC adjudication |
|
|
1 |
CPU reference |
Single run |
177.6 ms |
Apple M5 Max CPU |
2026-08-19 |
The current CUDA and MPS CoM rows pass their respective independent references.
The native row is the complete derived-product stage after exact moments exist;
it includes center/mean, alignment, Metal iDPC, and float-surface construction.
It does not include resident-cache traversal or source loading.
The CUDA/MPS cross-fixture timing rows are not compared numerically. The prior
same-fixture detector-bin-4 comparison remains a historical block for iDPC at
2.84e-5 maximum error. On the retained M5 Max hardware run, WebGPU DPC row
and column were byte exact. Optimized and zero-rotation iDPC had zero frozen
tolerance violations; optimized rotation had maximum absolute error
1.52587890625e-5 and maximum tolerance ratio 0.8993483035.
Single-sideband ptychography — quantem.gpu.SSB#
These are square scan-grid sizes, not detector dimensions.
Current 512x512 operation timing is separated from the size-support matrix.
Source loading, G(\mathbf k,\boldsymbol{\nu}) preparation, and UI paint are
excluded.
Platform |
Computer |
State |
Operation |
Detector plan |
BF policy |
Boundary |
Statistic |
Time |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Partial |
Complex object |
Native |
8,928 active |
Warm resident GPU |
p50 |
13.883 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Partial |
Exact phase |
Native |
8,928 active |
Warm resident GPU |
p50 |
32.035 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Partial |
Exact phase and loss |
Native |
8,928 active |
Warm resident GPU |
p50 |
32.335 ms |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Partial |
Exact phase and loss |
Explicit detector bin 2 to |
2,275 calibrated |
Single synchronized reconstruction |
Single run |
497.187 ms |
Apple M5 Max ( |
2026-08-19 |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Partial |
Complex object |
Native-detector exact BF columns |
9,074 logical / 2,459 executed |
Warm complete Hermitian cache |
p50 |
8.911 ms |
Apple M5 Max ( |
2026-08-19 |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Partial |
Exact phase-variance loss |
Native-detector exact BF columns |
9,074 logical / 2,459 executed |
Warm complete Hermitian cache |
p50 |
25.120 ms |
Apple M5 Max ( |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Refuted diagnostic |
Complex object |
Native |
3,418 active |
Readback-complete compute wall |
p50 |
32.5 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Refuted diagnostic |
Exact phase |
Native |
3,418 active |
Readback-complete compute wall |
p50 |
102.1 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
Refuted diagnostic |
Exact phase and loss |
Native |
3,418 active |
Readback-complete compute wall |
p50 |
189.4 ms |
Chrome 151, Apple M5 Max Metal-3 |
2026-08-19 |
CPU reference |
Portable CI runner |
Reference |
SSB |
Frozen adjudication only |
Frozen fixture |
Reference |
— |
Pending |
— |
— |
The WebGPU SSB values are retained diagnostic timings, not accepted scientific
performance. All 262,144 phase values differed from the frozen reference and
the wrapped maximum error was 0.0597773 rad against a 0.0002 rad gate.
That implementation remains refuted until phase parity is restored.
The raw detector-bin-2 MPS phase agrees with an independent CUDA reference to
1.2815e-6 wrapped radians maximum; loss differs by 7.45e-9. A prepared
BF-column companion from the same campaign is rejected because its stored
columns do not match the declared detector-bin coordinate grid. Its faster
timings are not published as scientific results.
Calibration is a separate operation:
Platform |
Computer |
Search |
Refinement |
Repetitions |
Statistic |
Time |
Result |
Device tested |
Date tested |
|---|---|---|---|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
Seeded Optuna TPE, 200 trials |
Nelder–Mead |
3 |
p50 |
11.168 s |
Byte-deterministic parameters, phase, object, and loss |
NVIDIA RTX PRO 6000 Blackwell Max-Q, GPU 1 |
2026-08-19 |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
Optuna TPE, 200 trials |
Nelder–Mead |
— |
— |
Pending |
Current compatible source not profiled |
— |
— |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
Seeded TPE, 200 trials |
Nelder–Mead |
3 |
p50 |
6.061 s |
Deterministic parameters and loss |
Apple M5 Max ( |
2026-08-19 |
WebGPU |
N/A (not implemented) |
— |
— |
— |
— |
— |
Unsupported |
— |
— |
CPU reference |
Portable CI runner |
— |
— |
— |
— |
— |
Reference only |
— |
— |
Levenberg–Marquardt is not implemented in any current SSB backend. An earlier CUDA atomic-objective calibration split into two different fitted minima under an identical seed; it is retained as a rejected experiment, not a benchmark row.
Platform |
Computer |
Scan grid |
Source kind |
BF policy |
State |
|---|---|---|---|---|---|
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
|
Fixed-size parity |
Frozen fixture |
Test |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
|
Fixed-size parity |
Frozen fixture |
Test |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
|
Native real acquisition |
Full active BF |
Partial physical evidence |
CUDA |
Linux CUDA workstation (dual 96 GB Blackwell GPUs) |
|
Fixed-size parity |
Frozen fixture |
Test |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
|
Resized/synthetic |
Fixed-size fixture |
Test |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
|
Resized/synthetic |
Fixed-size fixture |
Test |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
|
Native real acquisition |
Full active BF |
Partial physical evidence |
Python MPS |
MacBook Pro (M5 Max, 128 GB) |
|
Synthetic |
Fixed-size fixture |
Test |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
|
— |
— |
Not supported |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
|
— |
— |
Not supported |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
|
Native real acquisition |
Full active BF |
Partial physical evidence |
Native Swift/Metal |
MacBook Pro (M5 Max, 128 GB) |
|
— |
— |
Not supported |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
Real BF30 parity |
Radius 30 px |
Partial |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
Deterministic fixture |
Test BF |
Test |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
Real interaction |
Frozen phase reference |
Refuted |
WebGPU |
MacBook Pro (M5 Max, 128 GB) |
|
Real interaction |
Incomplete frozen reference |
Partial |
CPU reference |
Portable CI runner |
|
— |
— |
Not retained |
CPU reference |
Portable CI runner |
|
— |
— |
Not retained |
CPU reference |
Portable CI runner |
|
Independent adjudication |
Frozen fixture |
Ref |
CPU reference |
Portable CI runner |
|
— |
— |
Not retained |
Native Swift/Metal has a package-owned 512×512 implementation. Other native Swift scan sizes remain unsupported rather than inferred from CUDA/MPS. Untimed CUDA and MPS sizes retain fixed-size parity coverage. The WebGPU 512×512 phase result is explicitly refuted; its timing cannot be promoted until parity is restored.
See the SSB performance history for size-specific historical experiments. Current numerical rows remain in this dashboard and the verified-results ledger.
Cross-module platform map#
A one-row-per-platform map hides the computer, configuration, and evidence state, and previously made a refuted WebGPU SSB result look accepted. It is no longer maintained. Use the filterable atomic coverage matrix, where every row begins with Platform and Computer and keeps module, bin, dtype, cache state, and next gate separate.
A warm resident kernel, prepared source, first-process application load, and saved-result reopen answer different questions. See Benchmark methodology and Verified benchmark results before comparing them.
The public scientific array is always
where \(\mathbf R=(R_r,R_c)\) is the probe/scan coordinate and \(\mathbf k=(k_r,k_c)\) is the detector coordinate. Runtime-specific layout, tiling, fusion, and dispatch are private optimizations; shape, sampling, precision, calibration, and provenance are shared contracts.
Where an implementer starts#
Goal |
Read first |
Then inspect |
Acceptance evidence |
|---|---|---|---|
Change scientific meaning or add an operation |
Operation-specific equation, provenance schema, and independent reference |
||
Optimize CUDA |
The domain’s |
Real NVIDIA profile plus exact/frozen parity |
|
Optimize Python on Apple Silicon |
The domain’s |
Physical Apple device profile plus exact/frozen parity |
|
Build a native Apple client/library |
|
|
|
Optimize a browser client |
Domain WebGPU TypeScript/WGSL resources |
Real adapter, headed browser gate, and matching scientific output |
|
Deploy CUDA behind a process boundary |
Deployment, protocol, and admission pages |
Same array/provenance contract plus transport and capacity checks |
|
Add or review a benchmark |
Continuous profiling, parity, and the optimization ledger |
Date, revision, device, source plan, cache state, memory, wall boundary, and parity artifact |
Dashboard maintenance rule#
Update a dashboard row only after its detailed evidence row is complete. The detail remains authoritative and must record measurement date, exact source revision, physical device/runtime, source shape and dtype, cache state, crop/bin/load plan, benchmark definition, peak memory or swap where available, and numerical or hash parity. Keep an older result when the newer experiment changes any of those conditions; label both instead of silently replacing one. Keep every row atomic: a different bin, dtype path, fixture, cache state, or statistic is another row. Documentation tests enforce that the landing page stays timing-free and that current overview values remain owned by this dashboard.
Accepted and rejected experiments remain in the
optimization ledger, and the
machine-readable evidence fingerprints are in
performance/evidence_manifest.json.
The platform/module cadence, runner ownership, and known harness gaps are
machine-checked from benchmarks/profile_matrix.json; see
Continuous profiling.