Backends#
The canonical source-tree and cross-language test organization is documented
in Backend layout and parity contract.
Backend implementations remain private behind one domain-level scientific API;
folder moves require the machine-readable parity gates in
tests/parity/backend_matrix.json.
quantem.gpu supports three backend names:
Backend |
Purpose |
Notes |
|---|---|---|
|
NVIDIA GPU IO, decompression, reductions, and SSB reference paths |
Uses CuPy/CUDA kernels where available. |
|
Apple Silicon Metal/MLX paths |
Used for ANS-encoded loading, BF/DF/DPC images, and SSB fit/reconstruction/preview paths. |
|
Test/reference implementation |
Available only when explicitly requested by a test or reference comparison. |
Check the selected backend:
import quantem.gpu as qgpu
backend = qgpu.device.detect()
print(backend)
Request a backend explicitly when you need honest failure:
qgpu.device.resolve("cuda") # raises if CUDA is unavailable
qgpu.device.resolve("mps") # raises if MPS is unavailable
Use backend="auto" for normal scripts and backend="cuda" or
backend="mps" in parity/performance tests.
WebGPU is a browser runtime, not a Python device backend. Reusable browser
compute sources live beside their scientific domains and are packaged for
browser clients. SSB-specific WebGPU implementation files live under
ssb/webgpu. Browser performance claims must log a real
adapter; SwiftShader or another software adapter is only a smoke test.
CUDA service placement can select several devices for independent resident datasets. Their VRAM is not combined to make one dataset fit. WebGPU does not expose CUDA-style multi-GPU placement; the browser selects one adapter for a page.
Backend coverage#
CUDA and MPS are the native production backends. CPU is test/reference only and is never selected as a silent scientific fallback.
Status terms: Done means implemented with real-data parity and performance
evidence; Partial means source exists but the full signoff matrix is not
complete; Gap means the backend does not implement that capability yet.
Capability |
CUDA |
MPS |
WebGPU |
CPU |
Notes |
|---|---|---|---|---|---|
Device report and explicit selection |
Done |
Done |
NA |
Done |
WebGPU adapter selection is browser-side. |
HDF5 metadata, readiness, discovery |
Done |
Done |
Done |
Done |
Keep one shared API for all clients. |
Full HDF5 bitshuffle/LZ4 load/decompress |
Done |
Done |
Done |
Reference |
CUDA and MPS load complete acquisitions into ANS-encoded residents that keep native counts. CUDA and WebGPU retain native |
Scan-region selection |
Done |
Done |
Done |
Reference |
CUDA and MPS load the complete acquisition and return a rectangle with |
Exact detector bin |
Done |
NA |
Done |
Reference |
|
BF/DF/ADF resident kernels |
Done |
Done |
Done |
Reference |
CUDA RawKernel, MPS Metal, and WebGPU WGSL selected reducers are implemented for |
Dense DF/ADF strategy |
Done |
Done |
Done |
Reference |
Dense masks use cached |
CoM/DPC resident kernels |
Done |
Partial |
Done |
Reference |
Detector-bin-4 CUDA/MPS CoM passes the frozen gate. The public MPS native-detector interaction sidecar is detector-bin-2 and is not full-resolution parity. WebGPU row/col DPC has full no-bin headed signoff on real hardware. |
Cached detector/DPC products |
Done |
Done |
Product-first Done / cache-read Done |
Cache-read Done |
CUDA and MPS build the product cache from one encoded load, including exact |
iDPC |
Done |
Partial |
Done |
Reference |
Current CUDA/MPS detector-bin-4 iDPC exceeds the frozen |
Ptychographic SSB preview |
Done |
Done |
Partial |
Reference |
Python MPS and the separate native |
Ptychographic SSB fit/reconstruction |
Done |
Done |
Partial |
Not target |
Python MPS supports its current parity shapes. Native Swift/Metal supports exact 512×512 reconstruction, phase-variance loss, and deterministic 200-trial TPE plus Nelder–Mead fitting; other native scan sizes remain gaps. |
Native Browser FFT ( |
NA |
Done |
NA |
Reference |
Native Swift/Metal product for already-transferred 2D BF/ADF/custom images. 512×512 must stay inside 120 Hz when warm. Not a Python MPS path. |
GIF/MP4 movie rendering |
Done |
Done |
NA |
Fallback |
CUDA/NVENC and Metal/VideoToolbox paths live here; presentation controls remain client-owned. |
Browser source ownership |
Done |
Done |
Done |
NA |
Reusable TypeScript/WGSL source lives beside each scientific domain. |
The rule for new heavy work is: implement the compute or IO path in
quantem.gpu, then let clients call the shared contract.
Benchmark ownership#
This page owns capability and source-boundary status only. Numerical results are not copied here:
Implementation overview is the current human-facing speed, memory, feature, and parity dashboard.
Verified benchmark results is the authoritative provenance ledger.
Optimization ledger preserves accepted and rejected experiments.
Keeping the backend map timing-free prevents a prepared-index reopen, resident kernel, historical campaign, or application first product from drifting into a misleading source-load comparison.
Adding a backend kernel#
For agents and maintainers, a new optimized path is not complete until the source, tests, documentation, and measured evidence land together.
Kernel family |
CUDA source |
MPS source |
WebGPU source |
Required gate |
|---|---|---|---|---|
HDF5 bitshuffle/LZ4 decode |
|
|
|
Corrected-frame checksum parity and load-stage timing. |
BF/DF/ADF masked sums |
|
|
|
Exact integer product parity and first/warm interaction timing. |
CoM/DPC |
|
|
|
Row/col CoM and centered DPC parity within |
Display colormap/histogram/log/FFT |
None; CPU reference |
|
|
Exact uint8 RGBA and 256-bin counts for linear/signed-log float32 fixtures; FFT agreement within the stated float precision. |
SSB object, phase, loss |
|
|
|
Same complete BF disk, aberrations, float32/complex64 parity, and interactive redraw timing. |
Movie rendering |
|
|
NA |
Frame parity and encoded movie smoke tests. |