# Cross-backend parity

The same scientific claim must survive CUDA, Python MPS, native Swift/Metal,
WebGPU, and the explicit CPU reference where that capability applies.

## Layers

1. **Contract:** public imports, signatures, shapes, dtypes, coordinate order,
   provenance, and honest unsupported behavior.
2. **Synthetic numerical:** deterministic odd/rectangular shapes, partial edge
   bins, masks, nonfinite display inputs, and exact small references.
3. **Frozen cross-backend:** every backend reads the same fixture and writes the
   same versioned result bundle.
4. **Real data:** full source shape/dtype with source and output hashes and no
   unreported crop/bin/precision change.
5. **Physical end to end:** the consuming application uses an exact local
   package revision and reports wall time and memory on the target device.

Source-presence, compile, or software-adapter tests never replace hardware
parity.

## Exact and floating contracts

These operations are byte-exact when they operate on integer evidence:

- HDF5 bitshuffle/LZ4 decode;
- bad-pixel masking;
- scan and detector integer-sum binning, including partial edges;
- detector sums and selected diffraction values;
- histogram counts and integer RGBA lookup outputs; and
- cache payload identity and metadata.

CoM, DPC/iDPC, FFT, and SSB use frozen operation-specific metrics. Reports
include maximum and high-percentile error where meaningful. Tolerances are not
widened to make a new backend pass.

## Current real-data gate

The 2026-08-19 full-scan campaign used revision `8c47a466` and one
byte-identical `512x512x192x192` `uint16` compressed-HDF5 fixture on CUDA GPU 0,
Apple M5 Max, and Apple M5. No scan crop was used and detector bin was explicit.

At detector bin 4, mean diffraction and exact total/BF/DF are byte-identical
across the three machines. Apple-to-Apple CoM/iDPC is also byte-identical; CUDA
CoM differs by at most `1.91e-6` and passes the frozen `1e-5` gate. CUDA-versus-
MPS iDPC differs by `2.84e-5` and is blocked at the frozen `rtol=1e-5`,
`atol=1e-5` gate.

At detector bin 1, integer products are byte-identical, but the public MPS
interaction sidecar uses detector bin 2 and therefore changes native-detector
CoM/iDPC. A direct full-resolution Metal diagnostic passes CoM with `7.63e-6`
maximum error, while iDPC remains blocked at `7.58e-5`. Detector-bin-2 and
detector-bin-8 timings have no equivalent retained real-data product bundle and
remain timing-only. See [Verified benchmark results](results.md) for the full
matrix and device/runtime details.

Native Swift/Metal SSB at `e1da9bc` has its own independently frozen 512×512
phase reference. The complete-cache path passes at relative L2 `5.86952296e-5`
and maximum wrapped error `5.62884106e-6` rad; the zero-cache exact path passes
at relative L2 `2.30151898e-5` and maximum wrapped error `2.48712759e-6` rad.
Three seed-42 200-trial TPE plus Nelder–Mead fits returned identical parameters
and loss. These values do not adjudicate other scan sizes, raw-HDF5 preparation,
or the physical 8 GB gate.

## Result bundle

A cross-backend result bundle contains:

```text
manifest.json       source, revision, parameters, backend, device, provenance
arrays/             canonical binary arrays with shape/dtype metadata
metrics.json        exactness or floating error metrics
timing.json         stage and wall-clock measurements
environment.json    OS, runtime, driver, and dependencies
```

Goldens are recaptured only by an explicit command and are never generated by
the backend being adjudicated during an ordinary test run.

## Coverage map

The machine-readable capability matrix is
[`tests/parity/backend_matrix.json`](../../tests/parity/backend_matrix.json).
Its guard ensures that every scientific domain names coverage for CPU reference,
CUDA, MPS, native Swift/Metal, and WebGPU and that every named gate still exists.

Run the portable contract layer with:

```bash
PYTHONPATH=src python -m pytest -q \
  tests/infrastructure/test_backend_parity_manifest.py \
  tests/parity/test_products_parity.py \
  tests/parity/test_dpc_rotation_agreement.py \
  tests/parity/test_display_parity.py \
  tests/contracts/test_ssb_backend_contract.py
```

Native Apple parity additionally runs:

```bash
swift test
swift test -c release --filter MetalSSBKernelsTests
```

Real CUDA, MPS/Metal, and WebGPU gates run only on their qualified hardware and
retain their output artifacts.
