Choose a kernel runtime

Choose a kernel runtime#

For Python notebooks and scripts, begin with installation and the Python API guide. This section is for implementation work.

The scientific operation determines what must be computed; the runtime page explains how to implement it efficiently on a particular device.

You are implementing for

Start here

Build and memory model

Apple GPU from Python

Python MPS

Python adapters over MLX/PyObjC/Metal and unified memory

NVIDIA GPU

CUDA

Python/CuPy with CUDA kernels and dedicated VRAM

Native Apple client/library

Native Swift and Metal

SwiftPM products, Metal resources, and unified memory

Browser GPU

WebGPU

TypeScript/WGSL, browser security, and explicit GPU buffers

Native Android client/library

Android Vulkan

NDK C++/C ABI, Vulkan shaders, and admitted packed residency

Independent adjudication

CPU reference

deterministic small NumPy/reference implementations

Shared implementation shape#

Each runtime follows the same layered contract:

public scientific API
    ↓
device and operation dispatcher
    ↓
runtime adapter
    ↓
kernel, shader, or reference operation
    ↓
typed result and scientific provenance

The public API owns coordinates, shapes, dtypes, units, errors, and provenance. The dispatcher selects only an explicitly available implementation. The runtime adapter owns allocation, layout, compilation, queueing, and synchronization. Kernels may optimize those private details but must return the same scientific result.

Loaded data uses one vocabulary across runtimes:

Field

Values

Meaning

representation

encoded, paired, packed, dense

physical count layout; runtime support varies

dtype

scientific value type

the value range and arithmetic contract

residency

host or runtime device location

where the physical payload remains

storage_schema

versioned internal format

which decoder/profile produced it; Python metadata names it resident_profile

CUDA and Python MPS expose this through io.Dataset4dstemGPU, native Swift/Metal through Metal4DSTEMResidentReceipt, and WebGPU through LocalH5LoadResult. Vulkan load plans and packed-session admission retain their documented dense/packed contracts. Python CUDA/MPS acquisition loading defaults to ANS and CUDA also accepts paired; Python dense requires backend="cpu", and Python has no packed representation. A representation enum is not a promise that every runtime can load that layout. See the representation support matrix before assuming an operation accepts a representation on a given runtime.

Repository map by layer#

Runtime

Discovery/dispatch

IO/decode

Detector and DPC

Reconstruction/display

Primary tests

Python MPS

device/select.py, operation protocols

io/hdf5/mps, resident/mps

detector/counts.py, detector/mps, resident/mps

ssb/mps

test_mps_*, MPS sections of parity tests

CUDA

device/select.py, operation protocols

io/hdf5/cuda, resident/cuda

detector/cuda

ssb/cuda

test_cuda_*, test_realdata_parity.py

Swift/Metal

SwiftPM products in Package.swift

Native4DSTEMIO, Metal4DSTEMKernels

Metal4DSTEMKernels

MetalImageFFT, MetalDisplayKernels, MetalImageRuntime

native/swift/Tests

WebGPU

device/webgpu.ts and TypeScript adapters

io/hdf5/webgpu

detector/DPC WebGPU modules

SSB and display WebGPU modules

test_webgpu_*, browser hardware gates

CPU reference

explicit backend="cpu"

io/hdf5/cpu.py

NumPy paths in detector/DPC workflows

small independent reference fixtures

product/parity tests

Open the runtime page for exact source files, call path, memory behavior, profiling boundaries, build commands, and acceptance gates.

Before working in a platform folder, read the corresponding page under Scientific kernels. The platform may optimize layout, fusion, queueing, and transfers; it must preserve the operation’s (row, column) ≡ (r, c) contract and provenance.

backend="auto" is suitable for ordinary Python use. Tests and benchmarks select a runtime explicitly so missing hardware and unsupported paths fail honestly. Capability status is maintained in Backend coverage; current measurements live in Verified benchmark results.

Serving CUDA to another process is not a kernel implementation. See QuantEM.GPU Remote for local loopback and SSH-tunneled deployment.