Load detector acquisitions#
Reads a 4D-STEM acquisition onto CUDA or Apple Metal and returns an encoded
quantem.gpu.io.Dataset4dstemGPU, accepted directly by
Show4DSTEM. Public import:
from quantem.gpu.io import load
The QuantEM.GPU I/O guide is the authoritative description of supported sources, exactness, metadata, storage, device selection, and saving. For image files rather than scanned detector acquisitions, see Image and acquisition I/O.
Function reference#
- quantem.gpu.io.load(source: str | PathLike[str] | Sequence[str | PathLike[str]], *, dtype: str | type | dtype | None = None, backend: str = 'auto', representation: DataRepresentation | str | None = None, dataset_path: str | None = None, scan_shape: tuple[int, int] | None = None, scan_order: Literal['row-major', 'serpentine'] = 'row-major', scan_region: tuple[int, int, int, int] | Sequence[tuple[int, int, int, int]] | None = None, detector_region: tuple[int, int, int, int] | None = None, apply_mask: bool | None = None, hot_pixel_correction: str = 'median', auto_narrow: bool = True, stack: bool | None = None, device: int | str | None = None, verbose: bool = True) Dataset4dstemGPU | list[Dataset4dstemGPU]#
Load one or more 4D-STEM sources through an accelerated backend.
Supported original acquisitions default to bounded ANS ingestion on CUDA and MPS. Integer sources retain uint8/uint16; float32 sources retain IEEE bits at native detector geometry. Use
scan_shapefor headerless EMPAD RAW. CPU reference access is explicit, never a fallback. Generic HDF5 layouts use bounded storage reads where the direct compressed chunk decoder cannot apply. Saved copies retain calibration and provenance.Fractional intensity exports support explicit
dtype="scaled_uint16"on CUDA and Metal/MPS. They remain ANS encoded and print a measured conversion report. The older bit-packed float16 acquisition profile is rejected; omit dtype to preserve original float32 bits. Scaled codes restore their saved intensity units for detector queries.scan_regionanddetector_regionselect values before resident allocation. New scaled storage automatically calibrates bounded regions in one pass; saved files retain their recorded calibration, including legacy global scales. With explicit precision,sourcecan also be a GPU array or an object providingshape,dtype, andblocks()of ordered GPU frames. Each generated frame must appear once in row-major order.io.load("display_master.h5", dtype="scaled_uint16")is approximate; preserve the original float32 file for exact scientific analysis.Here
dtypeselects stored precision, not calculation precision or a compression codec."scaled_uint16"stores calibrated integer codes and reconstructs float32 intensities on the GPU; it does not recover discarded precision. Plain"uint16"is not calibrated scaled storage. Omitdtypewhen reopening a precision file to retain its recorded values and calibration."uint16_scaled"is not a supported alias.Complete native HDF5 acquisitions can be loaded together into compact encoded accelerator storage with
stack=False. Encoded is the default on CUDA and MPS. Each result remains an independent encoded owner, not a dense stacked array. Stored detector-mask pixels use GPU median replacement by default before encoding.All spatial arguments use
(row, col)order.representationselects how the complete logical data is retained. Ordinary HDF5 selects"encoded"automatically on CUDA and MPS. Dense GPU overrides fail before allocation. Useloaded.read(scan_region=...)for bounded tensor access.Self-contained ANS files default to
representation="encoded"and retain stored native counts. A list of compressed or independently calibrated files returns one resident acquisition per file whenstackis omitted. Preparing that list exposes one detector session without materializing a dense stack. Explicitstack=Trueis rejected for these sources.representation="paired"streams complete uint16 acquisitions (one path or a list, each returned as its own source) into the CUDA paired-count tANS resident layout, and a saved paired resident form reopens under the same name without decoding. CPU reference expansion requiresbackend="cpu", representation="dense". Unsupported conversions raise instead of silently loading HDF5, expanding densely, or using CPU.apply_mask=Nonezeroes stored dead pixels in the CPU reference. Compact HDF5 loads usehot_pixel_correction; saved compact sources retain the correction already recorded in their metadata. File format and compression are detected from contents, independently of resident representation. Nodecompression=argument is needed.backend="auto"selects CUDA or MPS and never selects CPU silently.hot_pixel_correction="median"is the default for an ordinary HDF5 acquisition loaded into resident ANS storage. Stored detector-mask pixels are replaced on the selected GPU by the integer local 3x3 median before encoding. Use"zero"for zero replacement or"none"to retain raw masked-pixel counts.- Parameters:
source – One master/data HDF5 path, a folder, a list of master paths, or a standalone QuantEM encoded file.
dtype – Omit (or
"native") to keep stored values."scaled_uint16"selects calibrated precision storage for float32 sources.representation –
"dense","encoded"or"paired". The authenticated storage schema selects the exact decoder within each representation. When omitted, ordinary HDF5 uses encoded CUDA/MPS storage; saved ANS sources retain their recorded representation. GPU dense overrides are rejected before loading; use boundedreadselections from the encoded acquisition. Unsupported source/representation/backend combinations raise rather than transform implicitly. Representation never changes scan coverage, detector coverage, binning, calibration, or scientific dtype.backend –
"auto","cuda","mps", or explicit reference"cpu".stack – Omit to retain compressed or independently calibrated acquisitions as a list, while stacking CPU reference arrays. Use
Falseto request independent acquisitions explicitly.scan_shape – Optional full scan shape as
(row, col).scan_region – Optional precision-storage bounds as
(row_start, row_stop, col_start, col_stop).detector_region – Optional precision-storage bounds as
(row_start, row_stop, col_start, col_stop).
- Returns:
Data stays backend-resident.
- Return type:
Dataset4dstemGPU or list[Dataset4dstemGPU]
Tip
load keeps native detector sampling and the source count dtype (uint32 counts
that fit in uint16 are stored as uint16). It has no detector-binning or
dtype-cast option: the acquisition stays ANS encoded on the
GPU, so a full-resolution scan already fits. Pass a list of master paths to get
one acquisition per file.
Backend (CUDA / Apple Silicon)#
load detects the native GPU automatically: an NVIDIA box loads onto CUDA
and a Mac loads onto Apple Metal (MPS). Pass backend="cuda" or
backend="mps" to require one, and device= to choose a CUDA device.
Scientific loading does not silently fall back to CPU.
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
loaded = load("scan_master.h5") # CUDA on a workstation, MPS on a MacBook
Show4DSTEM(loaded)
The returned acquisition is encoded, not a dense array. A 512 x 512 x 192 x 192 uint16 Arina scan, 18 GiB as a dense array, occupies about 0.1 to 2 GiB of GPU memory depending on its counts:
print(loaded.shape, loaded.dtype) # (512, 512, 192, 192) uint16
print(loaded.logical_bytes / 2**30) # 18.0, the dense size
print(loaded.resident_bytes / 2**30) # encoded size on the GPU
The same encoded storage serves a MacBook and a workstation; there is no separate preview or memory-reduction mode.
The same is one shell command - see the CLI:
quantem show4dstem scan_master.h5.
Several files#
Pass a list of masters to load each one as its own acquisition:
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
masters = [
"/data/session/file_001_master.h5",
"/data/session/file_002_master.h5",
"/data/session/file_003_master.h5",
]
acquisitions = load(masters) # a list of Dataset4dstemGPU, one per master
viewer = Show4DSTEM(acquisitions)
The acquisitions are not stacked into one 5D array. Show4DSTEM opens them as
a comparison grid with one shared detector ROI, labels each panel with its
source file name, and requires one scan and detector shape across the list.
For a folder that is still being written, use
Show4DSTEM.from_folder(...), which opens
after the first master and appends the rest.
Bounded reads for ROI workflows#
Reconstruction, denoise, and ROI workflows read a rectangular scan region from the loaded acquisition. The read decodes only that region into a Torch tensor on the acquisition’s GPU:
from quantem.gpu.io import load
with load("scan_master.h5") as loaded:
patch_t = loaded.read(scan_region=(160, 293, 234, 367)) # row_start, row_stop, col_start, col_stop
print(patch_t.shape)
# torch.Size([133, 133, 192, 192])
Stops are exclusive. detector_region=(row_start, row_stop, col_start, col_stop) restricts the detector pixels the same way, and basic indexing such
as loaded[100:164, 100:164] returns the same kind of tensor. A complete read
is allowed only when its dense tensor fits the accelerator’s working memory, so
read bounded regions instead of the whole scan.
For drift-corrected time-series work, compute the source scan box from the shared specimen ROI plus a small halo, read that patch, then apply your existing subpixel sampler in local patch coordinates. Do not save the sampled patch as a new raw acquisition; drift is scan-position metadata, and detector counts stay physically unchanged.
Detector products#
Virtual-detector images and the mean diffraction pattern are computed on the encoded storage without a dense copy:
from quantem.gpu import detector
mean_dp = detector.mean(loaded)
center, radius = detector.fit_probe(mean_dp)
bright_field = detector.bf(loaded, center=center, radius=radius)
session = detector.prepare(loaded) # repeated queries; a list of acquisitions also works
Releasing GPU memory#
Close an acquisition when the last consumer is done with it, or load it in a
with block:
loaded = load("scan_master.h5")
...
loaded.close()
with load("scan_master.h5") as loaded:
...
A viewer borrows the acquisition it shows, so close the viewer first.