Load, decode, and bin#

The load pipeline converts compressed detector evidence into a typed, accelerator-resident 4D-STEM representation without changing its scientific meaning.

from quantem.gpu import io

result = io.load("scan_master.h5")

print(result.shape, result.dtype)
print(result.representation, result.residency)
print(result.logical_bytes, result.resident_bytes)

The default retains native detector sampling and ANS residency for supported original acquisitions. No representation option is needed. Decode a selected region through indexing, such as result[10, 12]; do not expand the acquisition to select it. Scientific binning and cropping require a separately supported operation and are never introduced as an automatic resource policy.

Coordinate and layout contract#

The logical output is

\[ I[R_r,R_c,k_r,k_c], \qquad (\text{row},\text{column})\equiv(r,c). \]

In plain terms, (row, column) is (r, c) for both scan and detector axes.

Storage shards may flatten \((R_r,R_c)\) into a frame index, and a device layout may be detector-major or packed. Dataset4dstemGPU metadata must still report the logical scan and detector shapes, source/output dtype, crop/bin plan, and source identity.

Representation, dtype, and residency#

Keep these three axes separate:

Axis

Public values

Question answered

Representation

encoded, paired, dense in Python; native Swift/Metal and Vulkan also use packed

How are all logical counts encoded?

Dtype

uint8, uint16, uint32, and supported floating types

What scientific value type is exposed?

Residency

host, CUDA device, Apple unified/device memory, or WebGPU device

Where is the physical payload retained?

packed does not mean uint8, and dense does not imply host memory. The format profile and schema remain provenance fields rather than additional public representation names. This lets new codecs evolve without changing scientist-facing algorithms.

Dtype and memory contract#

The load path keeps four precision decisions separate:

  1. source dtype — the detector counts stored in HDF5, commonly uint16;

  2. working dtype — compact decode/staging precision, such as an audited lossless uint8 low-byte path;

  3. accumulation dtype — widened integer precision used for detector or scan sums; and

  4. output dtype — the value type delivered to downstream kernels and recorded in provenance.

uint16 is exact for native uint16 counts only while every correction and sum fits 0 through 65,535. Detector binning (native Swift/Metal, WebGPU, and the CUDA browse views) therefore widens accumulation and, when needed, the result to uint32. uint8 is exact only for a native uint8 source or after a complete identity-bound range audit proves every corrected count is at most 255. A saturating uint8 output without that proof, such as the WebGPU clip8 browse decode, is a browse transform and must retain its saturation count. Python io.load keeps stored counts and has no uint8 cast.

For output detector bin \(b\) and \(w\) resident bytes per value,

\[ B_{\mathrm{payload}} =N_{R_r}N_{R_c} \left\lceil\frac{N_{k_r}}{b}\right\rceil \left\lceil\frac{N_{k_c}}{b}\right\rceil w. \]

This is only the resident payload. Peak memory also includes live compressed bytes, decode/reduction scratch, staging or upload buffers, allocator reserve, products, and other active GPU users. Report predicted payload and measured peak separately; on Apple unified memory also report process footprint, pressure, and swap.

Pipeline stages#

discover/open -> metadata and index -> read spans -> decode
              -> bad-pixel correction -> dtype/bin/layout -> resident result

Profile these stages separately where the storage/runtime makes that possible:

  1. source discovery, file open, and metadata;

  2. index mapping and compressed-span planning;

  3. storage read or memory-map page-in;

  4. bitshuffle/LZ4 decode;

  5. bad-pixel handling and value-range audit;

  6. dtype conversion and exact detector/scan reduction;

  7. destination allocation and layout conversion;

  8. first usable scientific product; and

  9. final provenance/cache work.

On unified memory, storage page-in, decode, and device access may overlap. Do not invent a separate “upload” number when bytes are mapped without a copy.

Source-page control is one stage label, not a synonym for every kind of cold state. A controlled F_NOCACHE run may still reuse a sealed source audit, while a new index root and destination allocation have their own state. Report those facts independently so a prepared artifact is never mistaken for first-source work.

Count-preserving detector binning#

For detector bin factor \(b\), each output detector pixel is the exact sum of one \(b\times b\) source block:

\[ I_b[R_r,R_c,k'_r,k'_c] =\sum_{i=0}^{b-1}\sum_{j=0}^{b-1} I[R_r,R_c,bk'_r+i,bk'_c+j]. \]

This preserves the complete scan grid and total detector counts. It does not crop real space, interpolate detector values, average counts, or label the result as native detector resolution. Incomplete detector-edge blocks are summed over the source pixels that exist, using the same rule on every backend.

The efficient path performs this sum while decoded chunks are already on the accelerator. It writes the final resident dtype/layout directly instead of materializing both a full unbinned volume and a second binned copy. Correctness still requires widened accumulation, explicit overflow behavior, bad-pixel ordering, original/output detector shapes, selected bin, and the resource-policy reason in provenance.

compressed chunk
      ↓ GPU decode
decoded source counts ──► bad-pixel policy ──► exact b×b sum
                                                   ↓
                                  final resident binned counts

Optimization model#

The reusable fast path should:

  • align reads to source shards or compressed blocks;

  • keep file descriptors, indexes, pipelines, and masks prepared;

  • use double or triple buffering only within the measured memory plan;

  • overlap read/decode with reduction when command dependencies permit;

  • fuse dtype conversion and detector binning with decode when parity remains exact;

  • reuse pinned/shared/device buffers rather than allocate per batch;

  • avoid per-batch host synchronization and device-to-host copies; and

  • materialize only the requested resident layout or product.

Compressed size is not a memory estimate. Admission includes decoded bytes, decoder scratch, output buffers, reduction products, allocator reserve, and other active GPU users.

Resource-policy boundary#

The package estimates the cost of each exact plan. A client may choose an automatic detector bin for a memory-limited machine only when it records and presents:

  • requested and selected detector bin;

  • source and output detector shapes;

  • scan region and scan bin;

  • source, accumulation, and resident dtypes;

  • resident payload, predicted peak, measured peak, and the measurement source;

  • the reason the plan changed.

The resulting array is binned evidence, not native-resolution evidence.

Source map and gates#

Layer

Source

Public Python contract

src/quantem/gpu/io

CUDA decode

src/quantem/gpu/io/hdf5/cuda

Python MPS/Metal decode

src/quantem/gpu/io/hdf5/mps

CUDA and Python MPS ANS residents

src/quantem/gpu/resident/cuda and src/quantem/gpu/resident/mps

WebGPU local-file decode

src/quantem/gpu/io/hdf5/webgpu

Native IO and Metal load plans

native/swift/Sources/Native4DSTEMIO and Metal4DSTEMKernels

Independent reference

src/quantem/gpu/io/hdf5/cpu.py

Acceptance covers exact decoded counts, odd detector shapes, incomplete edge bins, bad pixels, dtype/overflow behavior, full source identity, memory-budget failure, and cold/warm/prepared timing. See Cross-backend parity.