Choose a kernel runtime#
For Python notebooks and scripts, begin with installation and the Python API guide. This section is for implementation work.
The scientific operation determines what must be computed; the runtime page explains how to implement it efficiently on a particular device.
You are implementing for |
Start here |
Build and memory model |
|---|---|---|
Apple GPU from Python |
Python adapters over MLX/PyObjC/Metal and unified memory |
|
NVIDIA GPU |
Python/CuPy with CUDA kernels and dedicated VRAM |
|
Native Apple client/library |
SwiftPM products, Metal resources, and unified memory |
|
Browser GPU |
TypeScript/WGSL, browser security, and explicit GPU buffers |
|
Native Android client/library |
NDK C++/C ABI, Vulkan shaders, and admitted packed residency |
|
Independent adjudication |
deterministic small NumPy/reference implementations |
Repository map by layer#
Runtime |
Discovery/dispatch |
IO/decode |
Detector and DPC |
Reconstruction/display |
Primary tests |
|---|---|---|---|---|---|
Python MPS |
|
|
|
|
|
CUDA |
|
|
|
|
|
Swift/Metal |
SwiftPM products in |
|
|
|
|
WebGPU |
|
|
detector/DPC WebGPU modules |
SSB and display WebGPU modules |
|
CPU reference |
explicit |
|
NumPy paths in detector/DPC workflows |
small independent reference fixtures |
product/parity tests |
Open the runtime page for exact source files, call path, memory behavior, profiling boundaries, build commands, and acceptance gates.
Before working in a platform folder, read the corresponding page under
Scientific kernels. The platform may optimize layout,
fusion, queueing, and transfers; it must preserve the operation’s
(row, column) ≡ (r, c) contract and provenance.
backend="auto" is suitable for ordinary Python use. Tests and benchmarks
select a runtime explicitly so missing hardware and unsupported paths fail
honestly. Capability status is maintained in Backend coverage;
current measurements live in
Verified benchmark results.
Serving CUDA to another process is not a kernel implementation. See QuantEM.GPU Remote for local loopback and SSH-tunneled deployment.