Paired-count resident layout#
An opt-in second resident layout for native count acquisitions, next to the
byte-rANS StreamedCounts default. Every original count is retained exactly;
the layout changes bytes and query kernels only. Nothing in the default load,
dense or ANS-file paths changes when this layout is not requested.
Operation pipeline#
Native
uint8/uint16counts arrive as complete 512-scan blocks (PairedCounts.append).io.load(..., representation="paired")produces them from original bitshuffle+LZ4 HDF5: reader threads read whole shards with direct I/O into page-locked staging and parse the LZ4 block headers there; one producer stream copies, LZ4-decodes and bitshuffles into a ring of rolling frame buffers; the consumer stream encodes and indexes each filled buffer and returns it through an event. Shard reads run ahead across acquisitions, so a series is bounded by the drive, not by the GPU.Each detector pixel’s 512-scan stream is coded as consecutive count pairs with a 1024-state tANS table from 32 Poisson pair models; streams that are empty, constant, sparse (few nonzero counts) or incompressible keep exact alternative forms. Stream offsets are grouped (one 32-bit base plus 32 relative 16-bit offsets).
Exact per-scan sums over radial-angular pixel groups (64-pixel leaves, 16-leaf roots) are bit-packed as the interaction index.
detector.prepareselectsPairedSeriesCompute. A detector mask (or its difference from the previous mask) is decomposed into index fields plus residual pixels; the residual streams are decoded by one thread per stream (a 64-bit bit reservoir refilled four bytes at a time from a prefetched window, once per three coded pairs) with a warp-wide packed reduction per 32 scans. Outputs are exact integer virtual images and native diffraction patterns, as for the default layout.masked_sum(..., block_stride=k)runs the plan, index and residual kernels over every k-th 512-scan block only (strideargument; work items per chunk shrink toceil(blocks / k)), so a viewer can preview a moving mask on all acquisitions at about 1/k of the cost; the baseline for incremental masks is kept per stride.PairedCounts.savewrites the resident arrays once;PairedCounts.load(orio.loadon the file) reopens them with direct I/O and no decode.Reconstruction consumers read native count blocks back from the resident form:
PairedCounts.decode_blocks(first, scans)decodes whole 512-scan blocks of one chunk, andPairedFeed(sources, amplitude=True)iterates the blocks of a series on a prefetch stream (depthbuffers ahead, events in both directions) so a joint time-series ptychography update never touches a dense copy of the data.
Axes are (scan_row, scan_col, detector_row, detector_col); equations use
\(I[R_r,R_c,k_r,k_c]\) with \(\mathbf R=(R_r,R_c)\) and \(\mathbf k=(k_r,k_c)\).
Contract#
item |
value |
|---|---|
query ABI |
|
input |
complete multiples of 512 scans, native |
H5 loader |
|
virtual image |
exact |
diffraction pattern |
native dtype, invalid pixels reported as zero |
malformed stream |
|
saved form |
|
application streaming |
|
queued queries |
|
Coding parameters (32 models, 1024 states, 2-byte stream header, sparse mode preferred unless the paired stream saves at least two bytes) are fixed by the ABI string; any change needs a new ABI.
Source map#
piece |
location |
|---|---|
kernels: tables, encode, compact, decode, decode_range, frame, plan, residual, polar fields, index pack/sum, offsets, planner weights |
|
|
|
|
|
planner hooks in the streamed base |
|
dispatch |
|
H5 streaming loader, saved-form reopen, |
|
representation selector and magic detection |
|
tests |
|
Verification#
The contract tests build synthetic sources with a wide literal row, a
saturated row and an invalid pixel at 17x17, 19x19 and 257x257 detectors,
compare every mask and frame with direct sums (including uint64 sums above
2**32), reject a reserved header bit and a shortened stream extent, and
reopen a saved form byte-identically, and iterate two sources through
PairedFeed checking every block and amplitude against the raw counts. The
loading tests write four-shard
Arina-style masters with save_compressed_arina_h5, stream one and two of them
through io.load(representation="paired"), compare every count, a detector
mask and a frame with direct sums, and reopen a saved form through io.load
byte-identically. Set QUANTEM_GPU_PAIRED_REFERENCE to a JSON file (path,
poses, reference_VI_sha256) to check frozen full-array virtual-image
digests on a native source through the public loader; on 2026-09-09 all six
digests of one native 512x512x192x192 acquisition matched.
Physical-device evidence (private study, 2026-09-09; one RTX PRO 6000
Blackwell, 69 complete 512x512x192x192 uint16 acquisitions resident): the
paired layout decoded arbitrary-center detector updates at 12.5 ms per batch of
69 full virtual images versus 23.9 ms for the byte-rANS layout on the same
counts, with 414 frozen full-array digests exact; resident bytes per source
1.293 GiB versus 1.307 GiB. Loading from the saved form reached the measured
drive ceiling. The grouped-refill decoder (2026-09-09, twelve sources, 356
recorded centre poses) executes 48 instead of 67 warp instructions per coded
pair and takes 0.72 instead of 1.00 ms per update, byte-identical
(docs/performance/data/paired-decoder-2026-09-09.json); on a viewer’s
device the residual decoder is 88 percent of query time and the plan kernel
10 percent.
Package loader, single run on the same device with the default
PairedLoader (69 complete 512x512x192x192 uint16 acquisitions offered from
one PCIe 4.0 NVMe drive whose direct-read ceiling is 5.5 GB/s; a memory-budget
admit callback stopped the series at 66):
measurement |
acquisitions |
value |
|---|---|---|
series load, original HDF5 to resident, wall (28.8 s once chunk tables are parsed from the staged image) |
66 |
29.3 s |
resident-ready wall per acquisition inside the pipeline, median |
66 |
1.50 s |
encode per acquisition, median |
66 |
0.197 s |
index per acquisition, median |
66 |
0.120 s |
resident bytes, sum |
66 |
93.97 GB |
device bytes in use after load (pool cache and staging included) |
66 |
101.6 GB |
|
66 |
38 ms |
first bright-field virtual image, all sources (first source digest exact) |
66 |
8.2 ms |
first diffraction pattern, all sources |
66 |
1.9 ms |
save one source as the paired resident form (write plus fsync) |
1 |
1.39 s |
reopen that form through |
1 |
0.53 s |
Shards are read with direct I/O, so the series time is bounded by the drive
(about 145 GB of original chunks in 29.3 s). The complete atomic record with
boundaries, cache states and the admission rule is
paired-loading-2026-09-09.json.
The last acquisitions of a full device fit only with a smaller loader
(PairedLoader(rolling_scans=512, rings=2)); that tail policy belongs to the
application.
Complete series through the package only (private owner script over
PairedLoader, PairedCounts.load and detector.prepare; 69 acquisitions of
512x512x192x192 uint16 on one device, single run):
measurement |
value |
|---|---|
original HDF5 to resident, 64 with the default loader then 5 with |
38.7 s |
resident bytes, sum (index 6.7 GiB of it) |
90.3 GiB |
peak device bytes in use during loading |
94.4 GiB |
|
53 ms |
frozen full-array virtual-image digests, six masks per source |
414 exact |
headless center gestures, 64 fresh masks, wall p50 / p95 |
14.1 / 16.3 ms |
headless scan gestures (native diffraction pattern), wall p50 |
1.3 ms |
save every source as the paired form, 39 on one drive and 30 on another |
192 s |
reopen every saved form, one reader per drive, to 69 resident |
9.6 s |
Record:
paired-capacity-2026-09-09.json.
Reopening carves every array of a file from one device allocation; with one
allocation per chunk array the driver’s allocation granularity cost about 7 GB
across the series and only 66 files fit.
Feeding a reconstruction from the resident form (three 512x512x192x192 sources,
PairedFeed(depth=2), consumer idle, same device shared with a desktop):
measurement |
per acquisition |
|---|---|
native counts, 512-scan blocks |
76 ms |
native counts, 8192-scan blocks |
90 ms |
counts plus float32 sqrt amplitude, 512-scan blocks |
119 ms |
counts plus float32 sqrt amplitude, 8192-scan blocks |
133 ms |
That is 127 G counts/s decoded into native frames, below the 48 ms per
262,144-position fused ptychography iteration only by a factor of about two, so
one decode per iteration hides behind the update when the two overlap on
separate streams. Record:
paired-feed-2026-09-09.json.
Device tested: RTX PRO 6000 Blackwell (CUDA). Date tested: 2026-09-09.