Load, decode, and bin#
The load pipeline converts compressed detector evidence into a typed, accelerator-resident 4D-STEM representation without changing its scientific meaning.
from quantem.gpu import io
result = io.load("scan_master.h5")
print(result.shape, result.dtype)
print(result.representation, result.residency)
print(result.logical_bytes, result.resident_bytes)
The default retains native detector sampling and ANS residency for supported
original acquisitions. No representation option is needed. Decode a selected
region through indexing, such as result[10, 12]; do not expand the acquisition
to select it.
Scientific binning and cropping require a separately supported operation and
are never introduced as an automatic resource policy.
Coordinate and layout contract#
The logical output is
In plain terms, (row, column) is (r, c) for both scan and detector axes.
Storage shards may flatten \((R_r,R_c)\) into a frame index, and a device layout
may be detector-major or packed. Dataset4dstemGPU metadata must still report the
logical scan and detector shapes, source/output dtype, crop/bin plan, and
source identity.
Representation, dtype, and residency#
Keep these three axes separate:
Axis |
Public values |
Question answered |
|---|---|---|
Representation |
|
How are all logical counts encoded? |
Dtype |
|
What scientific value type is exposed? |
Residency |
host, CUDA device, Apple unified/device memory, or WebGPU device |
Where is the physical payload retained? |
packed does not mean uint8, and dense does not imply host memory.
The format profile and schema remain provenance fields rather than additional
public representation names. This lets new codecs evolve without changing
scientist-facing algorithms.
Dtype and memory contract#
The load path keeps four precision decisions separate:
source dtype — the detector counts stored in HDF5, commonly
uint16;working dtype — compact decode/staging precision, such as an audited lossless
uint8low-byte path;accumulation dtype — widened integer precision used for detector or scan sums; and
output dtype — the value type delivered to downstream kernels and recorded in provenance.
uint16 is exact for native uint16 counts only while every correction and
sum fits 0 through 65,535. Detector binning (native Swift/Metal, WebGPU, and
the CUDA browse views) therefore widens accumulation and, when needed, the
result to uint32. uint8 is exact only for a native uint8 source or after a
complete identity-bound range audit proves every corrected count is at most 255.
A saturating uint8 output without that proof, such as the WebGPU clip8 browse
decode, is a browse transform and must retain its saturation count. Python
io.load keeps stored counts and has no uint8 cast.
For output detector bin \(b\) and \(w\) resident bytes per value,
This is only the resident payload. Peak memory also includes live compressed bytes, decode/reduction scratch, staging or upload buffers, allocator reserve, products, and other active GPU users. Report predicted payload and measured peak separately; on Apple unified memory also report process footprint, pressure, and swap.
Pipeline stages#
discover/open -> metadata and index -> read spans -> decode
-> bad-pixel correction -> dtype/bin/layout -> resident result
Profile these stages separately where the storage/runtime makes that possible:
source discovery, file open, and metadata;
index mapping and compressed-span planning;
storage read or memory-map page-in;
bitshuffle/LZ4 decode;
bad-pixel handling and value-range audit;
dtype conversion and exact detector/scan reduction;
destination allocation and layout conversion;
first usable scientific product; and
final provenance/cache work.
On unified memory, storage page-in, decode, and device access may overlap. Do not invent a separate “upload” number when bytes are mapped without a copy.
Source-page control is one stage label, not a synonym for every kind of cold
state. A controlled F_NOCACHE run may still reuse a sealed source audit, while
a new index root and destination allocation have their own state. Report those
facts independently so a prepared artifact is never mistaken for first-source
work.
Count-preserving detector binning#
For detector bin factor \(b\), each output detector pixel is the exact sum of one \(b\times b\) source block:
This preserves the complete scan grid and total detector counts. It does not crop real space, interpolate detector values, average counts, or label the result as native detector resolution. Incomplete detector-edge blocks are summed over the source pixels that exist, using the same rule on every backend.
The efficient path performs this sum while decoded chunks are already on the accelerator. It writes the final resident dtype/layout directly instead of materializing both a full unbinned volume and a second binned copy. Correctness still requires widened accumulation, explicit overflow behavior, bad-pixel ordering, original/output detector shapes, selected bin, and the resource-policy reason in provenance.
compressed chunk
↓ GPU decode
decoded source counts ──► bad-pixel policy ──► exact b×b sum
↓
final resident binned counts
Optimization model#
The reusable fast path should:
align reads to source shards or compressed blocks;
keep file descriptors, indexes, pipelines, and masks prepared;
use double or triple buffering only within the measured memory plan;
overlap read/decode with reduction when command dependencies permit;
fuse dtype conversion and detector binning with decode when parity remains exact;
reuse pinned/shared/device buffers rather than allocate per batch;
avoid per-batch host synchronization and device-to-host copies; and
materialize only the requested resident layout or product.
Compressed size is not a memory estimate. Admission includes decoded bytes, decoder scratch, output buffers, reduction products, allocator reserve, and other active GPU users.
Resource-policy boundary#
The package estimates the cost of each exact plan. A client may choose an automatic detector bin for a memory-limited machine only when it records and presents:
requested and selected detector bin;
source and output detector shapes;
scan region and scan bin;
source, accumulation, and resident dtypes;
resident payload, predicted peak, measured peak, and the measurement source;
the reason the plan changed.
The resulting array is binned evidence, not native-resolution evidence.
Source map and gates#
Layer |
Source |
|---|---|
Public Python contract |
|
CUDA decode |
|
Python MPS/Metal decode |
|
CUDA and Python MPS ANS residents |
|
WebGPU local-file decode |
|
Native IO and Metal load plans |
|
Independent reference |
|
Acceptance covers exact decoded counts, odd detector shapes, incomplete edge bins, bad pixels, dtype/overflow behavior, full source identity, memory-budget failure, and cold/warm/prepared timing. See Cross-backend parity.