Native Swift/Metal SSB migration#
This record freezes the extraction of reusable native SSB compute from the earlier iOS implementation into QuantEM.GPU. It does not migrate application UI, session state, navigation, plots, or cache-policy presentation.
Source lineage#
Role |
Repository state |
|---|---|
Original native implementation |
|
Original SSB sequence |
|
QuantEM.GPU base |
|
Extracted compute commit |
|
The original reusable compute was embedded in an application controller. The
migration moved the FFT, SSB reconstruction, exact objective, and optimizer
into a SwiftPM library while leaving all UIKit and application ownership behind.
MetalSSBKernels imports only Foundation and Metal.
Package boundary#
Public surface |
Contract |
|---|---|
|
Complete calibrated BF geometry in logical source order; row and column reciprocal arrays remain explicitly named |
|
|
|
Plane-major lossless |
|
Row-major 512×512 complex64 object and Fourier sum |
|
Exact full-logical-BF phase-variance objective |
|
Deterministic seeded TPE search, 200 trials by default, then Nelder–Mead |
|
Source/compute dtype, no crop, scan bin 1, logical/executed/proven-zero BF counts, cached/streamed BF counts, and exact cache bytes |
The engine retains all logical BF terms in normalization. The execution union may omit only BF terms proven to remain outside the aperture. It does not crop or bin scan positions, change precision, or fall back to CPU.
cacheBudgetBytes: nil requests the complete Hermitian cache. A finite budget
caches complete 32-BF batches and streams the remaining terms from the source
buffer with the same objective. Resource-policy choice and user-visible
explanation remain application responsibilities.
Real-reference acceptance#
The retained private fixture is identified publicly as
native-ssb-fullbf-512-u8-v2; paths and sample names are intentionally omitted.
Field |
Value |
|---|---|
Date and revision |
2026-08-19; |
Device |
Apple M5 Max ( |
Source |
9,074 plane-major BF images, 512×512, exact |
Source SHA-256 |
|
Reference |
Independent CUDA-formula 512×512 float32 phase |
Reference SHA-256 |
|
Scientific plan |
full 512×512 scan, scan bin 1, no scan crop, detector bin 1, float32/complex64 compute |
BF policy |
9,074 logical; 2,459 executed; 6,615 proven-zero; full logical normalization |
Cache state |
operating-system page cache warm; not a cold-source claim |
Complete cache#
Measurement |
Result |
|---|---|
File mapping |
4.703 ms |
Engine initialization |
74.342 ms |
Full cache preparation |
365.238 ms |
First reconstruction |
8.461 ms wall; 7.906 ms GPU |
Warm reconstruction |
8.911 ms p50; 9.416 ms p95/max, 7 repetitions |
First exact-loss call |
51.546 ms including one-time cache-layout change |
Warm exact loss |
25.120 ms p50; 25.516 ms p95/max, 7 repetitions |
200-trial TPE plus Nelder–Mead |
6.061212 s p50; 6.063051 s p95/max, 3 complete fits |
Fit repeatability |
identical fitted parameters and loss in all 3 fits, seed 42 |
Phase parity |
relative L2 |
Hermitian cache |
2,588,520,448 bytes |
Measured peak process footprint |
2,921,529,992 bytes; no process swaps |
The fused Hermitian exact-loss path replaces the generic cached objective. On the same prepared source it reduced the single loss measurement from about 144 ms to 25.120 ms p50 while retaining the independent phase gate. The generic streaming path remains the exact low-memory reference.
Zero cache#
Measurement |
Result |
|---|---|
Preparation |
0.001 ms |
First reconstruction |
312.677 ms wall |
Warm reconstruction |
145.178 ms p50; 147.070 ms p95/max, 7 repetitions |
Warm exact loss |
263.005 ms p50; 266.441 ms p95/max, 7 repetitions |
Phase parity |
relative L2 |
Measured peak process footprint |
310,510,216 bytes; no process swaps |
The two cache policies are not competing scientific modes. They produce the
same full-resolution result with slightly different float32 reduction order;
the focused gate limits cached-versus-streamed loss relative error to 5e-5.
Verification#
Check |
Result |
|---|---|
|
70 executed, 5 environment-qualified skips, 0 failures |
Focused debug and release SSB tests |
4/4 passed in each configuration |
Python SSB contract regression |
45 passed, 4 hardware-qualified skips |
Full-cache retained log |
SHA-256 |
Zero-cache retained log |
SHA-256 |
Limits and next gates#
The native engine currently supports a 512×512 scan. Other scan sizes remain unsupported, not inferred from the Python MPS/CUDA implementations.
The benchmark starts from an exact precomputed BF-column source. Raw HDF5 selection and BF-column construction are a separate IO/preparation stage.
No storage-cache reset was performed, so the preparation number is warm source evidence rather than cold-source wall time.
The complete cache fits the measured workstation but is not an 8 GB-device support claim. The zero-cache path proves a bounded exact alternative; a physical 8 GB app integration still needs its own admission and headed gate.
The package has no UI framework dependency. A consuming application must own controls, progress, memory-policy selection, and visualization.
The local commits are not pushed, released, or published by this migration.