Migration notes#
quantem.gpu exists to remove permanent duplicate accelerated code from
quantem.widget, quantem.live, and the legacy quantem.cuda name.
Current ownership#
quantem.gpu owns:
GPU IO and decompression.
Chunk assembly and load-to-device.
Device selection and backend errors.
Heavy BF/DF/DPC image compute.
SSB compute APIs.
quantem.widget owns:
anywidget UI.
Interaction state.
HTML/notebook export.
Display wrappers around arrays and reduced images from
quantem.gpu.
quantem.live calls quantem.gpu for product and SSB compute instead of
keeping second copies.
Dense, packed, and experimental status#
Explicit dense array algorithms and CPU references remain supported. Public GPU acquisition loading uses ANS; the Python package has no packed readers or packed load representation. Packed formats remain in the native Swift/Metal and Android/Vulkan code. Representation, location, and scientific dtype remain separate facts. Unsupported inputs never silently expand into a dense volume.
The following summarizes the representation contract,
not a new qualification registry. Exact gates remain in
tests/parity/backend_matrix.json; measured and pending performance remain in
benchmarks/profile_matrix.json.
Runtime |
Implemented entry points |
Remaining or experimental scope |
|---|---|---|
Python CUDA |
ANS acquisition loading; detector reductions and encoded products within their recorded contracts |
Physical CUDA acceptance remains separate from Apple tests. |
Python MPS |
ANS acquisition loading; detector reductions and encoded products |
Native Swift format qualification is separate from Python support. |
Native Swift/Metal |
Indexed dense and both packed profiles; source inspection, admission, authenticated loads, detector and prepared products |
App adoption and physical end-to-end qualification must use an exact package revision. |
WebGPU |
Dense HDF5 reader, encoded |
Experimental consumer integration. The held uint16 DPC/iDPC numerical candidate is not promoted; device-specific parity and presentation gates remain open. |
Android/Vulkan |
Native packed detector session, BF/DF/ADF, selected diffraction, bounded dense decode/staging |
Experimental. Compact headers support widths 0–8, expanded descriptors 0–16. No full dense-volume residency, shared SSB, or general 1024 FFT claim. |
CPU reference |
Explicit dense reference workflows and test decoders |
Not an automatic fallback or a public accelerated packed loader. |
These source changes do not establish full-file cold loading in 1–2 seconds or 120 presented scientific updates per second. Preparation, authentication, source reads, device residency, reconstruction, and actual presentation must be measured separately on the target device. Physical phone acceptance remains a consumer task after repinning; a host test or a resident kernel benchmark is not its substitute.
What the refactor changes#
The repository architecture defines one owner
per implementation. IO models, file formats, selection, and pinned staging have
separate modules. Scientific backend code lives in cuda/, mps/, and
webgpu/ directly under the package that owns the science; native Swift and
Android Vulkan sources live in native/swift and native/vulkan at the
repository root.
The import-only
compute/andbackends/modules and the earlier IO compatibility files are removed. Import from the package that owns the code.Retain the native library names in
native/vulkan/CMakeLists.txt.Remove unused private helpers only after checking callers. The reviewed cleanup removes obsolete loading, detector, and screening helpers.
Common imports no longer import CuPy or replace the caller’s pinned-memory allocator. CUDA allocation still uses CuPy’s configured allocator; the existing bounded host-registration pool remains in
device/cuda_runtime.py.io/load.pykeeps the load routes. The HDF5 decoders live inio/hdf5/(cpu.py,cuda/decode.py,mps/decode.py, and the WebGPU TypeScript).
Canonical representation names and receipt v3#
The Swift resident receipt is now quantem.gpu.4dstem-resident-receipt/v3.
representation is dense, packed, or encoded; the separate storage_encoding
field is removed. storage_schema, source/working dtype, geometry, hashes, and
byte counts retain the detailed scientific meaning. This is an explicit schema
change, not wire compatibility with v1 or v2. Existing sealed results keep their
original versions; do not rewrite old evidence to make it look like a new run.
Python callers use DataRepresentation.DENSE, .ENCODED, or .PAIRED
(quantem.gpu.io.representation). Swift clients use .dense, .packed, or
.encoded. The former ResidentStorageEncoding and
MPSResidentRepresentation type aliases and the
lossless_packed selector are removed. Update receipt parsers deliberately;
unknown schema versions must fail closed. Apple capability records are v4 and
publication records are v2. WebGPU uses the same three Swift names, and
Vulkan uses dense, packed, and ans; that does not qualify ANS loading or
kernels on those backends.
Save calls accept only format="arina" or format="quantem"; remove the old
HDF5 format aliases. File encodings themselves are unchanged. Ordinary native
HDF5 loads into encoded (ANS) storage by default on CUDA and MPS:
tilts = io.load(files, stack=False)
Each result is caller-owned and must be closed after its final consumer.
Next migration steps#
Pin each consumer to a reviewed package revision. Read the complete WebGPU source graph from the
webgpu/sources.jsonmanifest; build native clients from SwiftPM or the Vulkan CMake entry, without copying kernels.Adapt receipt parsers to v3 and Apple capability controls to v4. Keep unsupported operations unavailable rather than expanding or downcasting implicitly.
Verify original compressed HDF5 and prepared packed inputs separately, including 512 and 1024 scans where admitted, file A–B–A switching, cancellation, replacement release, and relaunch.
Run real detector translation and resizing, DF/ADF rings, DPC, colormaps, contrast, and FFT-off behavior in the actual app. Record scientific update cadence and presentation independently, with no hidden binning.
Close the held WebGPU numerical gates and missing backend operations before enabling them.
Native macOS Live4DSTEM calls the Swift package products instead of a local Python backend:
Client need |
Endpoint |
|---|---|
Browser FFT of BF/ADF/custom |
|
Histogram / contrast window |
|
HDF5/EMD catalog |
|
Decode / detector / CoM |
|
Remote raw 4D on a CUDA workstation |
Python |
Do not restore an app-local FFT, histogram, or HDF5 parser. Do not bundle
Python in the signed Mac app. Pin Live4DSTEM to an exact quantem.gpu
revision after this package is published; a local path override is only for
integration worktrees.
Release checks#
Before publishing an rc:
Run focused GPU parity tests.
Build wheel and sdist into a temporary directory.
Run
twine check.Inspect package contents for private data or generated reports.
Install from TestPyPI and verify:
import importlib.metadata as md import quantem.gpu assert md.version("quantem.gpu") == quantem.gpu.__version__
Do not regress#
Do not move GPU decompression back into widget.
Do not make SSB depend on anywidget.
Do not use
quantem.cudaas the public package name.Do not treat CPU fallback speed as acceptable for GPU workflows.
Do not use fast-mode SSB as parity evidence.
Do not copy
MetalImageFFTorNative4DSTEMIOsource into Live4DSTEM.Do not add a local Python FFT or HDF5 helper to the Mac app.
Python acquisition loading defaults to ANS#
io.load(path) now preserves complete native uint8/uint16 HDF5 counts in
lossless ANS GPU storage. Backend selection remains automatic. For multiple
acquisitions use io.load(paths, stack=False) to retain separate encoded owners.
Saved ANS and paired sources reopen their recorded encoded layouts. The Python
package no longer reads prepared-packed acquisitions; re-export their originals
as .qem.
Code requiring working tensors must request bounded loaded.read(...) regions.
Dense/packed GPU overrides, whole-acquisition tensor output, and the former
combined drift/stochastic loader interface are not supported. Explicit array
algorithms and tiny CPU reference loads remain separate APIs.
Ordinary HDF5 ANS ingestion is implemented for CUDA and MPS. Unsupported dtypes and backends raise with corrective guidance; there is no implicit dense or CPU fallback. CUDA and MPS precision loads keep encoded values resident and run detector queries on their owning accelerator. CUDA uses float64 intermediates where available; Metal uses deterministic float32/floating-pair reductions because Apple GPUs do not expose float64 arithmetic. Original ingestion uses bounded source blocks without retaining a complete decoded acquisition. Backend and format-specific staging are implementation details, not another user-selected representation.
The legacy dtype='u4' shortcut is rejected. Omit dtype to retain native
values in lossless ANS storage.
Explicit approximate precision for fractional intensities#
Keep a float32 archive, then explicitly choose a smaller working precision. The CUDA and MPS loaders retain scaled integer codes in ANS device storage and measure errors across every selected value on the accelerator. Loading does not change the source.
from quantem.gpu import io
io.save("merged_master.h5", merged, dtype="float32")
exact = io.load("merged_master.h5")
scaled = io.load("merged_master.h5", dtype="scaled_uint16")
The retired packed float16 loading profile is rejected; preserve the original
float32 archive instead. scaled_uint16 stores round((intensity - offset) / scale)
using recorded regional calibration. Returned patterns and reductions restore
code * scale + offset. These codes are not raw detector counts; code 65535
is a valid intensity. Plain uint16 keeps its existing whole-count meaning.
Backend choice and ANS encoding are automatic. Lossy precision is always explicit.
The loader reports source/working precision, intensity range, resident bytes, RMS and maximum absolute error, positive values becoming zero, overflow, clipping, and the number of values measured. Measurements use GPU reductions; no CPU codec or numerical fallback is used. Scaled uint16 may erase weak intensities despite a small RMS error. Preserve float32 for exact analysis.
region = io.load(
"merged_master.h5", dtype="scaled_uint16",
scan_region=(128, 256, 128, 256),
detector_region=(0, 192, 0, 192),
)
io.save("display_master.h5", region)
reopened = io.load("display_master.h5")
Bounds use (row_start, row_stop, col_start, col_stop). A list of scan regions
returns separately owned loaded objects. Global range measurement reads the
complete source in bounded GPU blocks; error measurement covers the selected
values. Reopened exports use saved scaling and label the original error report
as saved, rather than claiming a fresh comparison against the original source.
Inspect reopened.metadata["precision"] for the persisted report. close()
releases storage after the final consumer. Disk compression is GPU
bitshuffle/LZ4 for these HDF5 exports; ANS resident size and file size are different.
Export directly with io.save(..., dtype="scaled_uint16"). Conversion and writing use bounded GPU
blocks. A native 4D NPY source is also accepted by the precision loader.
Unsupported resampling, masks, and source dtypes fail explicitly. Nonfinite
sources are rejected by this approximate conversion path; preserve exact float32 instead.
Both CUDA and Metal implement scaled precision conversion, saving, and ANS reopening. For a file-backed source, the elapsed time includes reading the complete source; an already-resident MPS tensor uses the direct Metal path and avoids a host copy. The live widget consumes the loaded source without materializing a complete decoded array and exposes saved error details. The current release matrix still requires a dedicated minimum-memory laptop run before claiming a 24 GiB limit.