# Load, decode, and bin

The load pipeline converts compressed detector evidence into a typed,
accelerator-resident 4D-STEM representation without changing its scientific
meaning.

```python
from quantem.gpu import io

result = io.load("scan_master.h5")

print(result.shape, result.dtype)
print(result.representation, result.residency)
print(result.logical_bytes, result.resident_bytes)
```

The default retains native detector sampling and ANS residency for supported
original acquisitions. No representation option is needed. Decode a selected
region through indexing, such as `result[10, 12]`; do not expand the acquisition
to select it.
Scientific binning and cropping require a separately supported operation and
are never introduced as an automatic resource policy.

## Coordinate and layout contract

The logical output is

$$
I[R_r,R_c,k_r,k_c],
\qquad (\text{row},\text{column})\equiv(r,c).
$$

In plain terms, `(row, column)` is `(r, c)` for both scan and detector axes.

Storage shards may flatten $(R_r,R_c)$ into a frame index, and a device layout
may be detector-major or packed. `Dataset4dstemGPU` metadata must still report the
logical scan and detector shapes, source/output dtype, crop/bin plan, and
source identity.

## Representation, dtype, and residency

Keep these three axes separate:

| Axis | Public values | Question answered |
|---|---|---|
| Representation | `encoded`, `paired`, `dense` in Python; native Swift/Metal and Vulkan also use `packed` | How are all logical counts encoded? |
| Dtype | `uint8`, `uint16`, `uint32`, and supported floating types | What scientific value type is exposed? |
| Residency | host, CUDA device, Apple unified/device memory, or WebGPU device | Where is the physical payload retained? |

`packed` does not mean `uint8`, and `dense` does not imply host memory.
The format profile and schema remain provenance fields rather than additional
public representation names. This lets new codecs evolve without changing
scientist-facing algorithms.

## Dtype and memory contract

The load path keeps four precision decisions separate:

1. **source dtype** — the detector counts stored in HDF5, commonly `uint16`;
2. **working dtype** — compact decode/staging precision, such as an audited
   lossless `uint8` low-byte path;
3. **accumulation dtype** — widened integer precision used for detector or scan
   sums; and
4. **output dtype** — the value type delivered to downstream kernels and
   recorded in provenance.

`uint16` is exact for native `uint16` counts only while every correction and
sum fits 0 through 65,535. Detector binning (native Swift/Metal, WebGPU, and
the CUDA browse views) therefore widens accumulation and, when needed, the
result to `uint32`. `uint8` is exact only for a native `uint8` source or after a
complete identity-bound range audit proves every corrected count is at most 255.
A saturating `uint8` output without that proof, such as the WebGPU clip8 browse
decode, is a browse transform and must retain its saturation count. Python
`io.load` keeps stored counts and has no `uint8` cast.

For output detector bin $b$ and $w$ resident bytes per value,

$$
B_{\mathrm{payload}}
=N_{R_r}N_{R_c}
\left\lceil\frac{N_{k_r}}{b}\right\rceil
\left\lceil\frac{N_{k_c}}{b}\right\rceil w.
$$

This is only the resident payload. Peak memory also includes live compressed
bytes, decode/reduction scratch, staging or upload buffers, allocator reserve,
products, and other active GPU users. Report predicted payload and measured
peak separately; on Apple unified memory also report process footprint,
pressure, and swap.

## Pipeline stages

```text
discover/open -> metadata and index -> read spans -> decode
              -> bad-pixel correction -> dtype/bin/layout -> resident result
```

Profile these stages separately where the storage/runtime makes that possible:

1. source discovery, file open, and metadata;
2. index mapping and compressed-span planning;
3. storage read or memory-map page-in;
4. bitshuffle/LZ4 decode;
5. bad-pixel handling and value-range audit;
6. dtype conversion and exact detector/scan reduction;
7. destination allocation and layout conversion;
8. first usable scientific product; and
9. final provenance/cache work.

On unified memory, storage page-in, decode, and device access may overlap. Do
not invent a separate “upload” number when bytes are mapped without a copy.

Source-page control is one stage label, not a synonym for every kind of cold
state. A controlled `F_NOCACHE` run may still reuse a sealed source audit, while
a new index root and destination allocation have their own state. Report those
facts independently so a prepared artifact is never mistaken for first-source
work.

## Count-preserving detector binning

For detector bin factor $b$, each output detector pixel is the exact sum of one
$b\times b$ source block:

$$
I_b[R_r,R_c,k'_r,k'_c]
=\sum_{i=0}^{b-1}\sum_{j=0}^{b-1}
I[R_r,R_c,bk'_r+i,bk'_c+j].
$$

This preserves the complete scan grid and total detector counts. It does not
crop real space, interpolate detector values, average counts, or label the
result as native detector resolution. Incomplete detector-edge blocks are
summed over the source pixels that exist, using the same rule on every backend.

The efficient path performs this sum while decoded chunks are already on the
accelerator. It writes the final resident dtype/layout directly instead of
materializing both a full unbinned volume and a second binned copy. Correctness
still requires widened accumulation, explicit overflow behavior, bad-pixel
ordering, original/output detector shapes, selected bin, and the resource-policy
reason in provenance.

```text
compressed chunk
      ↓ GPU decode
decoded source counts ──► bad-pixel policy ──► exact b×b sum
                                                   ↓
                                  final resident binned counts
```

## Optimization model

The reusable fast path should:

- align reads to source shards or compressed blocks;
- keep file descriptors, indexes, pipelines, and masks prepared;
- use double or triple buffering only within the measured memory plan;
- overlap read/decode with reduction when command dependencies permit;
- fuse dtype conversion and detector binning with decode when parity remains
  exact;
- reuse pinned/shared/device buffers rather than allocate per batch;
- avoid per-batch host synchronization and device-to-host copies; and
- materialize only the requested resident layout or product.

Compressed size is not a memory estimate. Admission includes decoded bytes,
decoder scratch, output buffers, reduction products, allocator reserve, and
other active GPU users.

## Resource-policy boundary

The package estimates the cost of each exact plan. A client may choose an
automatic detector bin for a memory-limited machine only when it records and
presents:

- requested and selected detector bin;
- source and output detector shapes;
- scan region and scan bin;
- source, accumulation, and resident dtypes;
- resident payload, predicted peak, measured peak, and the measurement source;
- the reason the plan changed.

The resulting array is binned evidence, not native-resolution evidence.

## Source map and gates

| Layer | Source |
|---|---|
| Public Python contract | `src/quantem/gpu/io` |
| CUDA decode | `src/quantem/gpu/io/hdf5/cuda` |
| Python MPS/Metal decode | `src/quantem/gpu/io/hdf5/mps` |
| CUDA and Python MPS ANS residents | `src/quantem/gpu/resident/cuda` and `src/quantem/gpu/resident/mps` |
| WebGPU local-file decode | `src/quantem/gpu/io/hdf5/webgpu` |
| Native IO and Metal load plans | `native/swift/Sources/Native4DSTEMIO` and `Metal4DSTEMKernels` |
| Independent reference | `src/quantem/gpu/io/hdf5/cpu.py` |

Acceptance covers exact decoded counts, odd detector shapes, incomplete edge
bins, bad pixels, dtype/overflow behavior, full source identity, memory-budget
failure, and cold/warm/prepared timing. See
[Cross-backend parity](../performance/parity.md).
