Lossless Pack Format v1#
QuantEM.GPU Lossless Pack Format v1 stores an exact mask-applied integer 4D-STEM working array without constructing its dense logical tensor. The stable logical order is
(scan_row, scan_column, detector_row, detector_column)
All coordinates are zero-based (row, column), scan positions are C-order, and
detector pixels are C-order. Binning and crop are fixed to 1 and none. A
backend must reject conflicting metadata, unsupported widths, truncated
ranges, bad integrity hashes, or an output type that cannot represent the
requested exact reduction.
The public format has two exact encoding profiles. Their legacy binary revision numbers select the decoder and remain in existing files for compatibility; they are not separate public format generations.
Public format |
Encoding profile |
Legacy binary revision |
Scan tile |
Header representation |
|---|---|---|---|---|
Lossless Pack Format v1 |
exact |
1 |
128 |
one u8 width per pixel/tile |
Lossless Pack Format v1 |
exact |
3 |
32 |
per-pixel base, 32-tile checkpoints, and packed width nibbles |
The legacy index magic selects the encoding profile. A reader must not use the raw-LZ4 interpretation for a direct-bitpacked file or transfer performance evidence between the two profiles.
Authenticated detector exclusions always read as zero through the scientific
working-array API. The exact-uint16/LZ4 profile can retain those raw streams in
its payload and declare masked_detector_payload_policy as
retained_exactly_in_payload. The exact-uint8/bitpacked profile can instead
reconstruct excluded raw streams when its manifest carries producer-proven
constant uint16 values in
masked_detector_raw_values and authenticates the exact excluded-index sequence
with masked_detector_pixels_sha256. A result must say whether it was compared
with the raw source or with the mask-applied working array; raw sentinels never
enter masked scientific products.
Container user block#
The HDF5 file begins with this little-endian prelude. Integer offsets are file offsets, not HDF5 object offsets.
Field |
Type |
Required value or meaning |
|---|---|---|
magic |
8 bytes |
|
JSON bytes |
u32 |
UTF-8 manifest byte count |
JSON CRC-32 |
u32 |
IEEE CRC-32 of the exact manifest bytes |
binary offset |
u32 |
start of the binary index |
binary bytes |
u32 |
complete binary-index byte count |
The JSON region ends no later than the binary offset. Both ranges must be inside the file. A reader validates the CRC before interpreting the manifest.
The exact-uint16/LZ4 binary index begins with:
Field |
Type |
Required value or meaning |
|---|---|---|
legacy magic |
8 bytes |
|
shard count |
u32 |
positive |
payload chunk bytes |
u32 |
exactly 128 for the raw-LZ4 profile |
scan rows, scan columns |
2 x u32 |
positive logical scan shape |
detector rows, detector columns |
2 x u32 |
positive logical detector shape |
scans per shard |
u32 |
positive; the final 128-scan tile may be partial |
excluded count |
u32 |
no greater than detector pixel count |
excluded pixels |
repeated u32 |
unique row-major detector indices |
source identity |
32 bytes |
binary SHA-256 identity |
shard records |
repeated 96 bytes |
exactly |
The full scan plane must satisfy
scan_rows * scan_columns == shard_count * scans_per_shard
No trailing binary-index bytes are allowed. Payload, length, and width ranges must be disjoint and inside the file.
Each 96-byte shard record contains seven u64 values followed by two u32 values and a 32-byte digest:
payload_offset, payload_bytes,
lengths_offset, lengths_bytes,
widths_offset, widths_bytes,
decoded_bytes,
descriptor_count, chunk_count,
decoded_sha256
decoded_sha256 authenticates the exact bit-packed decoded payload before a
backend publishes the shard as resident.
Exact uint8/bitpacked binary index#
The direct-bitpacked profile uses legacy magic QGIX\0\0\0\x03 and replaces the
raw-LZ4 header’s payload
chunk size with three u32 values: reserved zero, scan_tile=32, and
header_encoding=1. Every shard must contain complete 32-scan tiles.
The 96-byte shard record is reused with strict direct-bitpacked meanings:
payload_offset,payload_bytes: directly addressable little-endian u32 payload;decoded_bytesmust equalpayload_bytes;lengths_offset,lengths_bytes,chunk_count: all zero because there is no raw-LZ4 envelope;widths_offset,widths_bytes: compact u32 headers rather than raw-LZ4 u8 widths;descriptor_count: compact header word count; anddecoded_sha256: SHA-256 of the exact direct payload bytes.
For T = scans_per_shard / 32, each detector pixel stores
ceil(T / 32) checkpoint words followed by ceil(T / 8) words holding eight
four-bit widths each. Header word zero is the detector pixel’s payload base.
Subsequent checkpoint words equal the cumulative widths at tiles 32, 64, and
so on. Widths are 0 through 8; unused tail nibbles are zero. The final pixel’s
base plus cumulative width must end exactly at payload_bytes / 4. Every
excluded detector pixel has all-zero widths and consumes no payload.
The binary index authenticates each direct-bitpacked payload but does not hash the compact headers. Qualification therefore also requires an externally frozen whole-file SHA-256, or a future schema revision that embeds header digests. A filesystem path or adjacent sidecar by itself is not an integrity identity.
Required exact-uint16/LZ4 JSON agreement#
The exact-uint16/LZ4 profile retains the legacy manifest schema
quantem.gpu.packed-detector-h5/v1. These fields must
agree exactly with the binary index:
source_shapesource_identity_sha256shard_countscans_per_shardpayload_chunk_bytesmasked_detector_pixels
The source dtype is uint16, scan_bin and detector_bin are 1, crop is
null, and status is complete. An exact-uint16/LZ4 reader accepts
working_dtype equal to
uint8 or uint16, but it derives values from descriptor widths rather than
casting to that label.
Early raw-LZ4 files can declare working_dtype: uint8 while retaining 16-bit width
metadata for excluded raw-sentinel columns. Those columns have authenticated
zero payloads. Compatibility is limited to widths above eight whose detector
pixel is in the authenticated exclusion list. A width above eight for any
nonexcluded pixel conflicts with the legacy manifest and fails closed.
The general exact-uint16 producer declares working_dtype: uint16 and retains
every source sample, including samples at excluded detector pixels, in its
packed payload. It declares
masked_detector_payload_policy: retained_exactly_in_payload, which permits
the raw reconstruction API to read those retained samples. The mask-applied API
still returns zero for excluded detector pixels, so detector products never
consume excluded values. An exact-uint16/LZ4 source with exclusions but
without this explicit
policy is not admissible as a portable raw reconstruction source.
Optional source-bound detector calibration#
A prepared source may carry one reusable detector calibration in the JSON manifest. Readers validate it before exposing it to a product UI:
{
"detector_calibration": {
"schema": "quantem.gpu.detector-calibration/v1",
"source_identity_sha256": "<same identity as the compact source>",
"detector_center_px": [97.26, 95.67],
"bright_field_radius_px": 41.0,
"dpc_rotation_degrees": 176.98,
"dpc_component_order_exchanged": false,
"method": "mean-diffraction-half-maximum"
}
}
Detector coordinates are always [row, column]. Center coordinates and the
bright-field radius must be finite and in range. The DPC rotation and component
order are optional, but must either both be present or both be absent. Binding
the calibration to source_identity_sha256 prevents a calibration copied from
another file from being silently accepted. Existing exact-uint16/LZ4 sources
without this
optional object remain valid and can be calibrated once after residency.
Required exact-uint8/bitpacked JSON agreement#
The exact-uint8/bitpacked profile retains the legacy manifest schema
quantem.gpu.packed-detector-h5/v3. It must bind the
binary index to the complete source and working representation with:
source_identity_sha256,source_raw_logical_sha256,source_shape, andsource_dtype: uint16;working_dtype: uint8,prepared_uint8_sha256, and the exactworking_value_definition;detector_mask_sha256and the ordered row-majormasked_detector_pixelslist;masked_detector_pixels_sha256, the SHA-256 of the exact ordered little-endian u32masked_detector_pixelssequence, whenever any detector stream is omitted;masked_detector_raw_values, an exact uint16 list aligned one-to-one with that pixel list, whenever any detector stream is omitted;scan_bin: 1,detector_bin: 1,crop: null,scan_tile: 32, and exactshard_count; andpayload_codec: direct-bitpacked-u32withstatus: complete.
Legacy files that omit these provenance fields are not portable
exact-uint8/bitpacked acceptance fixtures even when an adjacent audit can fill
the gaps. Produce a new immutable
enriched and sealed copy or rebuild the artifact; do not weaken readers based
on local path conventions.
detector_mask_sha256 preserves a distinct provenance identity: it hashes the
complete little-endian uint32 calibration/source mask, including its original
mask values and zero entries. It must not be reinterpreted as the digest of the
sparse index list. masked_detector_pixels_sha256 binds that sparse list to the
binary index and equals
sha256(pack_little_endian_u32(masked_detector_pixels)).
The exact-uint8/bitpacked working representation intentionally returns zero
at excluded pixels.
masked_detector_raw_values is ordered by the same strictly increasing
row-major pixel indices as masked_detector_pixels; each item is an integer in
the exact uint16 range 0 through 65535. Before publishing the file, the producer
must independently prove that every omitted sample across every scan and shard
equals its recorded value. The field records the result of that proof; it is not
a license for a reader to infer a sentinel from width or mask metadata.
The native Lossless Pack Format v1 producer implements the source-authenticated inspect-plan-produce lifecycle, bounded shard construction, atomic publication, and versioned receipt for this contract. Its current CPU reference backend is explicit in provenance and is not presented as GPU compression.
The Swift loader records the raw access mode in
MetalCompactH5Metadata.rawAccessMode. It is exact_exclusion_constants only
when masked_detector_raw_values and masked_detector_pixels_sha256 both
validate against the binary-index exclusion list. Scientific detector sums
continue to exclude those pixels.
Pre-release direct-bitpacked files with nonempty masked_detector_pixels but
missing either
masked_detector_pixels_sha256 or masked_detector_raw_values remain readable
only as explicitly nonportable mask-applied candidates. The Swift loader marks
them mask_applied_only_legacy, and
Metal4DSTEMResidentCapabilities.compact rejects that mode. They cannot
qualify as portable raw-lossless sources, even if an adjacent audit happens to
contain the missing values. Produce a new immutable enriched and
whole-file-sealed copy; do not mutate the earlier file.
Exact-uint16/LZ4 descriptor and payload layout#
One descriptor covers 128 consecutive scans for one detector pixel. Descriptor order is detector-pixel major, then scan-tile major:
descriptor_index = detector_pixel * tile_count + scan / 128
tile_count = ceil(scans_per_shard / 128)
The file stores one unsigned width byte per descriptor. The resident descriptor is reconstructed as one u32:
descriptor = (payload_word_offset << 5) | width
Widths are 0 through 16. Offset uses the high 27 bits. A tile consumes exactly
4 * width u32 words because it contains 128 * width bits. Consecutive
offsets must meet exactly, and the last descriptor must end at
decoded_bytes / 4. A width of zero represents 128 zeros without payload.
For local scan s, detector pixel p, and descriptor width w:
bit = (s % 128) * w
word = payload_word_offset + bit / 32
shift = bit % 32
value = low w bits spanning word and word + 1 when required
Bits are least-significant-bit first inside little-endian u32 words. There is no implicit saturation, low-byte cast, transpose, crop, or bin.
The decoded bitstream is split into independent 128-byte raw-LZ4 blocks. One
stored u8 length per block encodes compressed_bytes - 1. Lengths must cover
the compressed payload exactly. Chunk count is ceil(decoded_bytes / 128);
only the last decoded chunk can be shorter, and its size remains u32-aligned.
Mask and exact interaction semantics#
The exclusion list is separate from user detector geometry. Every selected diffraction pattern returns zero at an excluded pixel. A virtual detector first normalizes its row-major zero-or-one mask by clearing excluded pixels.
Virtual-detector sums use exact u32 output only when the maximum representable sum implied by the selected per-pixel widths fits u32. Otherwise the current adapter rejects the request and asks for a narrower mask or a future u64 path. After a full rebase, detector movement can use signed membership deltas. The next scan map is published only after the whole GPU command succeeds.
FFT work is independent of detector reduction. An FFT-off interaction must report zero FFT dispatches. A backend may not recompute FFTs during detector movement and still label that measurement FFT-off.
Resident tilt-series updates#
When several exact sources are already resident on one device, a backend may
apply the same detector mask to all of them in one queue submission. Every
source keeps a distinct persistent scan-map output; publication occurs only
after the shared submission completes, so a failed update cannot expose a
partially updated tilt series. Duplicate source objects, mixed devices, and
different detector (row, column) geometries are rejected.
This batch operation reduces command-submission and presentation overhead. It
does not merge, downcast, bin, crop, or otherwise shrink the source residents.
Memory qualification must therefore report the sum of the measured resident
allocations for the selected sources. The Swift/Metal entry point is
MetalCompactH5ResidentSource.updateVirtualDetectors.
Bounded resident lifecycle#
A conforming exact-uint16/LZ4 loader processes one shard at a time:
validate both metadata copies and all file ranges;
read one compressed payload plus its length and width metadata;
reconstruct and validate exact descriptors;
decode one shard into bounded staging;
verify the shard’s decoded SHA-256;
copy the decoded payload and descriptors to backend-private storage;
release transient shard storage; and
publish the complete source only after every shard succeeds.
For a failed or cancelled load, no partially resident source is returned. The
19.3 GB logical uint16 cube for a 512 x 512 x 192 x 192 source is never a
temporary or resident allocation in this path.
A direct-bitpacked loader instead validates the compact headers, reads/authenticates the direct payload and its headers, and copies those compact buffers to backend-private storage. It does not run LZ4 or expand descriptors for all pixel/tile pairs. Publication still occurs only after every shard and the whole-file header integrity identity pass.
Benchmarks report metadata, whole-file integrity, source read, host header validation, direct-payload integrity, descriptor preparation, private upload, GPU header validation or decode, resident-ready, first selected diffraction, and first detector-ready phases separately. A run is not described as cold unless storage and operating-system cache state were deliberately controlled and recorded.
Adapter status and gates#
Adapter |
Contract code |
Current qualification gate |
|---|---|---|
Swift and Metal |
|
physical Metal load, all decoded shard hashes, selected diffraction, complete detector-map oracles, bounded allocation, FFT-off interaction |
The Python, CUDA and WebGPU readers of this container were removed in October
2026; io.load does not open it. Reopen the original acquisition and save a
.qem copy for Python workflows.