I/O API#

Start with From acquisition to images or the README examples for individual patterns, scan patches, detector crops, and acquisition series.

quantem.gpu.io has four public operations:

from quantem.gpu import io

files = io.discover("/data/session")
with io.load(files[0]) as data:
    pattern = data[10, 12]
    metadata = data.metadata
    io.save("copy.qem", data)

For ordinary work, load once and read metadata from the returned acquisition. inspect is optional: use it for header-only checks, acquisition-readiness polling, or metadata access without an accelerator.

Metadata parsing may run on the host, but detector decoding and compression do not silently fall back to CPU. backend="auto" selects CUDA or MPS and raises with a corrective message when neither accelerated backend is available. The explicit backend="cpu" path exists for reference and parity tests.

discover#

Find candidate HDF5 masters before inspecting or loading them:

masters = io.discover(
    "/data/session",
    pattern="*_master.h5",
    recursive=True,
    scan_shape=(512, 512),
)

The optional scan shape uses the public (row, col) convention and filters by frame count without decoding detector pixels.

inspect#

Read headers and external-link metadata without loading the 4D array:

report = io.inspect("scan_master.h5", scan_shape=(512, 512))
if not report.ready:
    raise RuntimeError(f"{report.reason} Next step: {report.action}")

The report includes the frame count, detector (row, col) shape, dtype, source layout, and a source signature suitable for acquisition-readiness polling. For .qem, inspection rejects an incomplete file length but does not read and authenticate the entire payload. io.load verifies payload checksums before exposing the resident measurements.

load#

Array indexing#

with io.load("acquisition.qem") as data:
    pattern = data[10, 12]
    region = data[8:12, 10:16, 64:128, 64:128]
    metadata = data.metadata

io.load returns a Dataset4dstemGPU acquisition handle. QuantEM.GPU owns its storage, decoding and metadata; QuantEM’s core data classes are unchanged.

Attribute or expression

Meaning

data.shape

Scan row, scan column, detector row, detector column sizes

data.ndim

Number of logical axes

data.size

Total number of logical detector values, not bytes

len(data)

Size of the first scan axis

data.dtype

Stored scientific dtype

data.sampling

Per-axis spacing; None where unknown

data.units

Per-axis physical units; None where unknown

data.origin

First-pixel coordinates; None where unknown

data.metadata

Calibration, provenance, and storage metadata

data.logical_bytes

Size of the complete uncompressed array

data.resident_bytes

Size reported by the resident storage owner

Indexing returns an ordinary PyTorch tensor on the source GPU. Install PyTorch for this indexing interface. The calibration properties are read-only tuples in axis order: scan row, scan column, detector row, detector column. Scan lengths use angstrom; detector spacing retains its recorded angular or reciprocal-length unit. Missing calibration is never invented. Selected Torch tensors do not carry the owner’s calibration automatically; pass the appropriate sampling and units explicitly when adding scale bars. A single scalar selection returns a scalar tensor. Integers remove axes; slices and one ellipsis support negative indices and steps. Index arrays, Boolean masks, and new axes are not supported. This selects measurements without interpolation. Strides may decode a bounding region. Request only the working region that fits the GPU memory budget.

Iterating over data yields scan rows, as array iteration does; it is not a way to retrieve storage fields. Use .data only when an advanced operation needs the encoded storage object. np.asarray(data) on encoded storage is rejected to prevent an accidental whole-acquisition host copy. If a host array is needed, explicitly select a bounded region first: data[10, 12].cpu().numpy(). Use data[10, 12] directly for PyTorch operations. For several acquisitions, select the owner first: series[1][10, 12]. The explicit read(scan_region=..., detector_region=...) operation is also available for region-based pipeline code.

Notebooks and scripts#

In a notebook, load with one assignment and keep the owner available across cells:

data = io.load("acquisition.qem")

Explore with detector.bf(data) or array indexing such as data[10, 12]. Run data.close() after the last use, including any viewer using its buffers. Close an old owner before replacing it by rerunning a loading cell.

In scripts and batch jobs, prefer automatic cleanup, including on exceptions:

with io.load("acquisition.qem") as data:
    pattern = data[0, 0]

Neither spelling changes storage, precision or backend selection. Do not use an operation handle or viewer after closing its acquisition.

Inspect the loaded acquisition#

Every supported source uses the same entry point. It returns Dataset4dstemGPU, which keeps backend-native data and its scientific/storage metadata together:

loaded = io.load("scan-lossless.h5", backend="auto")

print(loaded.shape)
print(loaded.dtype)
print(loaded.representation)
print(loaded.residency)
print(loaded.logical_bytes, loaded.resident_bytes)

For ordinary acquisition use, omit representation, compression, and backend. Supported originals become ANS-resident on the selected CUDA or MPS device; saved .qem files retain their declared encoded layout. Unknown formats are rejected, not silently expanded or relabeled as ANS.

Load original arrays into ANS and save a .qem copy#

NumPy arrays, EMPAD-G1 RAW/XML and calibrated EMPAD2 float32 exports now use the same encoded-device workflow as supported HDF5 sources:

from quantem.gpu import detector, io

with io.load("scan.npy", backend="mps") as loaded:  # or backend="cuda"
    dp = loaded[10, 20]
    mean_dp = detector.mean(loaded)
    bright = detector.bf(loaded)
    io.save("scan.qem", loaded)

dp is a Torch tensor on the source GPU. mean_dp and bright are reduced NumPy images; only those small products are transferred to the host. The complete acquisition stays ANS-encoded. The original-array ingestion path uses at most 32 MiB per input window, with separate bounded encoder scratch. Saving copies encoded bytes, original metadata and normalized scientific fields without expanding the cube.

Original source

Encoded loading

NumPy .npy

Four axes, uint8/uint16 or float32; wider integer counts require an exact range audit

EMPAD-G1 .raw / .xml

Float32 records; 130×128 storage, 128×128 detector

EMPAD2 .xml

Calibrated 128×128 float32 exports, not encoded sensor words

DigitalMicrograph .dm3 / .dm4

Native uint8/uint16 or float32; one calibrated four-axis image

NCEM EMD .emd

Four-axis arrays; dataset_path selects among multiple acquisitions

HDF5

4D datasets or flattened 3D frames, including contiguous and gzip layouts

Float32 ANS retains the source detector geometry, including rectangular detectors, without cropping or rounding to integer counts. A single float32 frame must fit the 32 MiB working-window limit; larger detectors use fewer frames per window. Headerless EMPAD RAW needs scan_shape=(rows, columns). XML calibration is retained. Raw EMPAD2 detector words still require matching sensor calibration and a qualified decoder; this does not establish EMPAD-G3 or arbitrary EMD support. The native bitshuffle/LZ4 HDF5 layout retains its direct GPU decoder. Other supported HDF5 layouts use bounded storage-library reads before GPU ANS encoding.

NCEM EMD coordinate vectors follow the EMD specification. Regular scan sampling is normalized to Å and reciprocal sampling to Å⁻¹ (angular sampling stays in mrad). Original coordinate values, labels and units are retained; nonuniform coordinates and unknown units are not guessed. The reader retains coordinate vectors up to 4,096 values and records omitted larger vectors explicitly. Arrays use scan-row, scan-column, detector-row, detector-column order; arbitrary Velox event/image layouts are not implied. Use io.inspect(path, dataset_path="experiment/acquisition/data") and the same dataset_path in io.load when an EMD contains multiple acquisitions.

GPU NumPy/EMPAD loading rejects a dense residency override. Explicit backend="cpu", representation="dense" remains available for reference access, not as an automatic fallback. Unsupported dtypes/layouts raise an actionable error before allocating a resident cube.

NumPy simulations stored as int32, such as those in the SrTiO3 dislocation dataset, are audited in bounded windows before allocation. With auto_narrow=True (the default), nonnegative counts up to 65,535 use uint8 or uint16 ANS without changing any value. The original dtype, complete count range, and exact-narrowing provenance survive .qem export. Negative or larger values are rejected rather than clipped. EMPAD RAW loading finds a sibling XML that names the RAW file, retaining its scan dimensions, instrument fields, and independently indexed virtual-detector regions. If multiple XML files name the same RAW, open the intended XML explicitly. The archive’s separate para.txt is not automatically interpreted: keep it with the original acquisition and explicitly record any calibration taken from it. Its three-dimensional potential.npy is a simulated object, not a four-dimensional detector acquisition.

The TCMEP dataset contains prepared /dp HDF5 stacks and MATLAB 7.3 /cbed simulation arrays. Keep each prepared stack beside its params_backup.mat and, when supplied, data_position.hdf5. The Python reader validates the recorded raster positions or inclusive crop bounds rather than guessing a square scan. It retains the source parameters, normalizes voltage to kV, convergence semi-angle to mrad, and diffraction sampling to Å⁻¹. Object sampling dx is not scan sampling; position units that are not documented in the companion remain unspecified.

from quantem.gpu import detector, io

with io.load("data_roi0_Ndp128_dp.hdf5", backend="mps") as loaded:
    pattern = loaded[0, 0]
    io.save("acquisition.qem", loaded)

Use backend="cuda" for an NVIDIA device. Gzip HDF5 uses bounded host reads followed by GPU ANS encoding; this is not GPU gzip decompression. Float64 input is accepted only when auto_narrow=True and a complete bounded bitwise float64 → float32 → float64 audit proves every value unchanged. Otherwise it fails before resident allocation. Exact narrowing and the original dtype are recorded in .qem provenance; arbitrary float64 support is not implied. Reconstruction objects, probes and result TIFFs are not diffraction acquisitions. These Python paths do not certify the native Swift reader, which still limits float acquisition geometry to 128 × 128 and does not read these gzip stacks.

Stream complete H5 counts into CUDA encoded residency#

from quantem.gpu import io, detector

loaded = io.load("scan_master.h5", backend="cuda")
pattern = loaded[0, 0]

This default CUDA path streams bounded chunks of a complete uint8/uint16 H5 acquisition, preserves every stored count, and builds exact spatial sums while those chunks are available. The library’s default H5 representation on accelerator backends is encoded. Existing encoded files keep their original buffers. Keep acquisitions separately with io.load(paths, stack=False) and select an owner for each detector or reconstruction operation.

The resident’s resident_profile, physical_resident_bytes, index_bytes and load_timings metadata describe the actual loaded representation. The runtime H5 resident and the integer .qem codec share the runtime-column-rans-spatial-v2 profile, so native streamed residents are saved as .qem copies with the workflow below without re-encoding. Complete-series 120 Hz throughput is not established by the bounded CUDA parity tests.

Open native DigitalMicrograph camera counts#

Install the dm extra (pip install "quantem.gpu[dm]") to read DM3/DM4 metadata. Open a complete calibrated 4D diffraction image directly:

from quantem.gpu import detector, io

loaded = io.load("STEM SI.dm4", backend="cuda")
pattern = loaded[0, 0]
bright_field = detector.bf(loaded, center=(431.5, 431.5), radius=126)
# Close after all operations and viewers finish using the acquisition.
loaded.close()

The reader selects the unique 4D image and excludes embedded thumbnails and survey images. Native uint8/uint16 counts stream through bounded pinned staging into lossless CUDA ANS residency with exact spatial sums. Scan tails need not be multiples of 512. No crop, binning, clipping, detector masking, or intensity normalization is applied. Geometry and calibration use (row, col) order; metadata retains axis units, sampling, pixel origins and microscope voltage. Unsupported axis layouts and ambiguous multiple 4D images raise actionable errors. backend="cpu", representation="dense" explicitly opens a read-only memory map for reference access.

Save this encoded CUDA resident once to reopen it without re-encoding counts or rebuilding spatial indexes:

loaded = io.load("STEM SI.dm4", backend="cuda")
io.save("STEM SI.qem", loaded, format="quantem", backend="cuda")
loaded.close()
reopened = io.load("STEM SI.qem", backend="cuda")

This writes a QEMDATA1 (quantem.qem) copy of the exact ANS bytes, spatial indexes, detector validity and calibration. Reopening verifies the header and every 64 MiB block with SHA-256 while uploading through two bounded pinned buffers. It does not require the original DM file. Writes are atomic and reject existing destinations. Saved copies are detected by magic regardless of extension; the only accepted extension is .qem. DM selection/conversion options remain unsupported. Native uint8/uint16 DM counts also have an MPS encoded loading path; DM4 tests alone do not qualify all DM3 variants. Load differently shaped acquisitions separately; a list can use stack=False to return independent residents.

Reopen float32 .qem on CUDA or MPS#

For float32-bit-lanes-rans-v1 files, use the same public API on either accelerator:

from quantem.gpu import detector, io

with io.load("measurements.qem", backend="cuda") as loaded:  # or "mps"
    pattern = loaded[0, 0]
    mean_pattern = detector.mean(loaded)
    io.save("measurements-copy.qem", loaded)

The acquisition remains encoded on-device. The selected pattern is a PyTorch tensor; pattern.cpu().numpy() copies that selection to the host. The mean pattern is returned as NumPy. There is no CPU scientific fallback or complete dense-cube allocation. Copying retains original IEEE float bits, calibration, source documents and the saved background recipe. The output must not already exist. See the float codec contract for supported geometry, reduction precision and memory limits. This does not extend the collection converter below to raw float32 HDF5 inputs.

Converting a collection from the command line#

quantem-gpu convert /data/arina/session --dry     # GPU payload-size estimate; no copies written
quantem-gpu convert /data/arina/session           # writes name.qem beside each name_master.h5
quantem-gpu convert /data/arina/session --out /archive/session

Each published copy keeps master-file fields and attributes (units included) under its HDF5 path, embeds the master file itself so that long tables such as the flatfield survive (quantem.gpu.io.convert.restore_master writes it back byte for byte), records the name, size and SHA-256 of its source files, and fills the scientific metadata (source format, accelerating voltage, dwell time) from the Arina master. An Arina master records neither the probe semi-angle nor the scan step; when the acquisition’s folder has a session dataset.yaml, its microscope voltage and semi-angle and, through the files entry whose master is this file, the scan step of its magnification become calibration overrides whose evidence is the session file’s SHA-256 (quantem.gpu.io.convert.session_calibration). Readers then report them as the effective semiangle_mrad, voltage_kV and scan_sampling_A, so the copy needs no sidecar. The session’s values override what the master records (an Arina’s photon_energy field can hold a stale voltage). Only the fields used are attached (dataset.json), not the session notes. The command line names the calibrated fields for each copy. After writing, every value is compared with the detector files read through h5py before the copy is published. Flagged pixels are preserved and checked too: conversion disables display-time hot-pixel correction. CUDA uint32 files are stored as uint16 only after every stored value is shown to fit, with the original dtype retained in metadata. Larger values are refused, even at flagged pixels. float32 acquisitions are not supported by this collection command (other .qem writers support them). Use --backend cuda or --backend mps to select a device backend; the Metal collection loader currently accepts uint8/uint16, not uint32. An acquisition whose copy would be larger than its HDF5 files is reported and left as HDF5. Oversized masters that cannot be embedded are not converted, so long metadata tables cannot be silently lost. Source files are never modified or removed; existing copies are never overwritten. --no-verify explicitly skips the comparison. Dry-run size is the encoded payload estimate, not final file size, and still performs GPU loading and encoding.

Load an existing saved copy directly into native Metal#

The native Swift reader accepts the same quantem.qem files written by the Python MPS/CUDA and native exporters:

let snapshot = try NativeANSSnapshot(url: qemURL)
let source = try MetalRuntimeANSResidentSource.load(
  snapshot: snapshot, device: device, maximumAdditionalBytes: budget
)
let pattern = try source.extractRawDiffraction(scanRow: 0, scanColumn: 0)

It validates typed section bounds and checksums, streams bounded file ranges into private Metal buffers, and preserves native uint8/uint16 counts without a dense allocation. The path is independent of scan shape, including 512×512, 1024×1024, and non-square scans. This is an encoded-file reopen path; the macOS original HDF5 loader still performs bitshuffle/LZ4 decode on cold HDF5 input. The same handoff has been checked on bounded real 4D-STEM HDF5 slices; see the local experiment record for the retained parity evidence. Qualification of this saved-file reader does not establish cold original-HDF5 latency or application scheduling. Measure the selected original-file ingestion path separately from .qem reopen.

Stream complete uint16 H5 counts into the paired CUDA layout#

This is an advanced CUDA integration path with reusable native output buffers. For ordinary acquisitions, use io.load(path) and the automatic detector functions without a session or representation option.

from quantem.gpu import io, detector

series = io.load(masters, backend="cuda", representation="paired",
                 dtype="native", apply_mask=False)
session = detector.prepare(series)
images = session.masked_sum(detector_mask, output="native")
preview = session.masked_sum(detector_mask, output="native", out=images, block_stride=4)   # every 4th scan row, 1/4 of the time

representation="paired" accepts one master or a list. Each acquisition becomes its own exact PairedCounts source; a list is loaded by one streaming pipeline whose shard reads run ahead across file boundaries, so the series proceeds at the drive’s rate. The layout needs complete uint16 acquisitions whose scan count is a multiple of 512. Per-source load_timings report the direct-read, header-parse, encode and index seconds separately from the resident-ready wall time.

Save the resident arrays once and reopen them without decoding:

series[0].data.save("acquisition.paired")
reopened = io.load("acquisition.paired", backend="cuda", dtype="native", apply_mask=False)

The saved form starts with the fixed QGPUPAIR magic, so io.load selects "paired" for it automatically; asking for another representation on that file raises. Applications that admit acquisitions against a memory budget can drive quantem.gpu.io.PairedLoader.load_many(paths, admit=...) directly and stop the series while earlier files are still streaming.

Representation#

See Count representations for per-backend operation support, exactness, and ownership. Dense array algorithms are retained, but dense is not a public GPU acquisition loading mode.

representation describes how the complete logical array is retained. It has the following layout names; the loader admits only backend-appropriate choices:

Representation

Meaning

"dense"

Every logical value occupies its ordinary dense array element

"encoded"

Exact integer counts remain entropy-coded with the tables needed for decoding

"paired"

Exact integer counts in the CUDA paired-count tANS layout with a polar interaction index; explicit for original HDF5, detected for saved paired resident forms

These are the only representation names. The authenticated storage schema selects the precise decoder within a representation; users do not select an internal codec profile through this argument.

Native Swift/Metal can reopen a saved .qem copy directly, with a bounded physical geometry gate. That gate does not qualify complete-series loading, peak memory, or interactive throughput. The explicit CPU reference can decode encoded data to dense. Accelerated HDF5-to-ANS ingestion is implemented on CUDA and MPS. Native float32 coverage remains narrower than Python coverage. Do not infer support for every source/representation/backend combination from the selector names.

Representation is independent of dtype and residency. An encoded uint8 source and an encoded uint16 source have the same representation but different scientific dtypes. CUDA device memory, Apple unified memory, and host memory are residency locations, not representations.

The shortest call is source-native:

loaded = io.load("scan-lossless.h5", backend="auto")

Ordinary HDF5 uses ANS residency on accelerator backends. A saved .qem source remains encoded. Dense GPU loading is rejected. Read only the working region needed by the calculation:

data = io.load("scan_master.h5")
pattern = data[0, 0]

Tiny explicit CPU references may use backend="cpu", representation="dense". They are not a production loading fallback.

Selection#

Keep the acquisition compressed and request bounded working regions:

data = io.load("scan_master.h5")
crop = data[32:160, 48:176, :192, :192]

Indices follow (scan_row, scan_column, detector_row, detector_column) and slice stops are excluded. This selection does not bin or alter stored samples. Additional transformations belong to the requested working array, not an implicit loading policy.

Dtype selection#

Preserve original counts or float32 measurements by omitting dtype:

exact = io.load("scan_master.h5")

Selector

Meaning

Scientific boundary

dtype="native" or None

Preserve the backend-native source dtype

Preferred when the source precision must remain unchanged

dtype="scaled_uint16"

Explicit calibrated approximate intensity storage

Quantization is reported; not a lossless substitute for float32

The ordinary ANS count loader rejects dtype narrowing (u8, u16, auto) rather than silently changing scientific values. The retired bit-packed float16 acquisition profile is also rejected; reopen its original float32 source and save a lossless .qem copy.

Native uint8 input and a uint16 input converted to uint8 are different provenance. A lossless conversion requires a retained source identity, bad-pixel policy, maximum count, and pixelsAbove255 == 0. Otherwise retain the saturation count and label the result browse-only. Reconstruction workflows should retain the raw-count precision required by their objective. In particular, values above 255 cannot be preserved in an unsigned 8-bit array.

The resident payload is not peak memory. Record the requested, source, working, accumulation, and output dtypes; original/output shapes; bin/crop; payload bytes; predicted peak; measured process/accelerator peak; pressure or swap; and the resource-policy reason. See the dtype and peak-memory dashboard.

For a series, retain separate encoded acquisitions:

acquisitions = io.load(master_paths, stack=False)
# Close each owner after its last queued scientific operation.

io.load has no joint-series sampling, drift, or whole-acquisition tensor output. Request bounded patterns from each owner and pass explicit probe coordinates to the reconstruction workflow.

save#

Save backend-resident arrays without routing through a host reference writer:

saved = io.save(
    "processed_master.h5",
    data,
    backend="auto",
    dtype="u16",
    metadata={"scan_sampling_A": 0.264},
)
saved.wait()

backend="auto" infers CUDA from a CuPy array and MPS from an MPS tensor. A NumPy array requires backend="cpu" explicitly; this makes reference/test writes visible rather than accidental.

The default file contract remains an Arina-style master with external data files and lossless bitshuffle/LZ4 storage for integer detector counts. save always returns a completion handle. With the default wait=True, the handle is already complete; with wait=False, call saved.wait() before using the output.

File format, compression, and resident representation#

These are independent decisions, not different names for the same setting:

Option

Decision

Current examples

format on save

File layout

"arina", "quantem"

compression on save

Lossless file encoding

"bitshuffle_lz4" for Arina; "ans" for QuantEM

representation on load

In-memory count layout

GPU: "encoded", explicit CUDA "paired"; CPU reference: "dense"

For example, save and reopen an ANS-resident acquisition:

# Step 1. Save a copy of a complete encoded CUDA or Metal resident.
# Nothing is re-encoded; the resident's exact bytes and indexes are stored.
io.save("experiment.qem", resident)

# Step 2. Load the copy and select the representation used by GPU operations.
with io.load("experiment.qem") as data:
    print(data.shape, data.dtype, data.representation)

Loading detects the decoder from file contents. There is no decompression= argument, and changing a filename extension does not change the encoding. compression="auto" selects bitshuffle/LZ4 for Arina and ANS for QuantEM. Only format="arina" and format="quantem" are accepted: removed format aliases raise instead of redirecting the call. Incompatible format/compression pairs also raise. ANS is not an implemented HDF5 filter here. This API change does not change the bytes of the supported file formats.

The standalone writer is transactional and never overwrites an existing file. Source/working shape, dtype, and calibration are preserved. Saving and reopening an encoded acquisition is distinct from low-level representation conversion. Use io.inspect for supported-container metadata and readiness; successful inspection is not payload-integrity or full-volume performance qualification.

EMPAD correction provenance#

The reference exporter rejects an explicitly already-background-corrected float array until that correction state is qualified across readers. Keep that array and its metadata, or export the original measurements. Decoding does not perform an additional background subtraction. Original XML is retained; storage footer words are not detector measurements. See QEM interoperability for independent validation and the metadata mapping for which source fields are normalized or merely retained.