Virtual-Image Kernel Checklist#
Target: BF/DF/ADF, CoM/DPC, and arbitrary detector ROI dragging should use a custom GPU kernel on every production backend. Do not use detector binning, scan cropping, or CPU fallback as evidence for this checklist unless the row says so.
Why This Matters#
The future large browse target is 1024x1024x192x192 uint8, which is 36.0 GiB
resident. Old CuPy/Torch gather paths allocate large transient selected-pixel
slabs during every drag. A BF disk at radius 30 reads about 2,828 detector
pixels per scan position; a dense DF mask can touch most of the 36,864 detector
pixels. That is why the standard must be one custom kernel per product over
resident data, plus a dense-mask strategy such as total - complement.
Expected kernel paths#
The CUDA 1024x1024x192x192 uint8 target uses:
Product |
Expected path |
|---|---|
BF |
|
ADF |
|
DF |
|
CoM/DPC |
|
The MPS 1024x1024x192x192 uint8 target uses the same formulation:
Product |
Expected path |
|---|---|
BF |
|
ADF |
|
DF |
|
CoM/DPC |
|
The Show4DSTEM WebGPU browser runtime is widget-bundled, but the reusable source now belongs beside its scientific domain. The parity/performance standard is the same:
Product |
Expected path |
|---|---|
BF |
|
ADF |
|
DF |
|
CoM/DPC |
|
iDPC |
|
Backend Checklist#
Backend path |
Current status |
Required tests |
Performance gate |
|---|---|---|---|
Show4DSTEM Python CUDA, resident CuPy |
Implemented in |
Exact parity vs old CuPy selected-pixel sum for BF/ADF/DF and old CuPy CoM for DPC; widget smoke must report |
512x512x192x192 uint16 no-bin BF/ADF/DF/DPC faster than old widget path; 1024x1024x192x192 uint8 shape probe must pass before real-data allocation tests. |
Public |
Implemented through |
Exact parity for |
Same or faster than old CuPy gather path, with lower transient memory. |
Show4DSTEM MPS chunk-backed data |
Implemented for uint8/uint16 through |
Mac runtime parity vs NumPy/reference on BF/ADF/DF and CoM/DPC; widget smoke must report |
512 no-bin interaction should use the Metal path or fast sidecar; 1024 uint8 requires an explicit memory policy before real allocation. |
Show4DSTEM WebGPU browser |
Implemented in canonical domain-owned WebGPU sources. BF/DF/ADF uses |
Source contract test plus headed Chrome test with a real adapter, not SwiftShader; BF/ADF/DF, CoM, DPC, and iDPC parity against NumPy/Python reference. |
Drag path should keep VI, DPC, and iDPC GPU-resident where the display pipeline can accept GPU buffers; widget model-byte shims remain for current anywidget compatibility. |
Multi-tilt/series compare grid |
CUDA path tested for seven 512 panels at detector bin 2; full no-bin panels are one-at-a-time unless sharded. |
Compare-grid BF/ADF/DF parity and timing for 7 panels; verify per-panel backend path. |
Refresh all visible panels without falling back to per-panel Torch gather. |
Done For This Branch#
CUDA RawKernel selected-pixel reducer for resident CuPy
uint8anduint16.CUDA selected reducers use warp shuffle instead of repeated block-wide shared reductions.
CUDA dense DF uses cached integer
total - complementand fuses complement reduction, subtraction, and final float32 output in one kernel.CUDA total-count maps use a custom row reducer instead of CuPy’s generic
sum(axis=1)path.CUDA DPC/CoM uses a fused raw kernel that accumulates total intensity, detector-row moment, and detector-column moment in one pass over each diffraction pattern. The full-detector CoM field is cached per backend, so repeated DPC/iDPC requests do not reread the resident 4D block.
Shape-only support check covers
1024x1024x192x192 uint8for CUDA and MPS.MPS Metal
uint8masked-sum, detector-sum, mean-DP, bin-sidecar, radial-cache, and CoM kernels are present.MPS dense dark-field masks use the cached
total - complementpath, matching the CUDA dense-mask strategy.WebGPU source ownership is split across the relevant scientific domains, with the Show4DSTEM engine and ShowPtycho SSB engine copied as canonical source package data. Widget build/export syncs these sources before bundling. BF/DF/ADF has a GPU-resident buffer path; DPC row/col now uses a WGSL CoM reducer, WGSL mean reducer, one-ULP mean-side correction, and WGSL centered-component pass. Browser iDPC uses paired DPC buffers and a dual-real FFT before the Poisson integration.
Local full 512x512x192x192 no-bin WebGPU DPC/iDPC browser signoff on a real NVIDIA Blackwell adapter:
corrected-frame load parity: passed
DPC row/col max abs error:
7.63e-6iDPC mean abs error:
4.70e-6; max abs error:3.05e-5from float32 FFT orderDPC row/DPC col/iDPC display medians:
14.9/13.2/13.2 msDPC row/col/iDPC recompute medians:
13.7/19.3/22.7 msidle RAF:
60 FPSlocal-file timing reruns use
--require-local-profileso the browser URL fallback cannot be recorded as a local-file benchmark
Local full 512x512x192x192 and true crop-256 WebGPU detector-bin local-H5 signoff on a real NVIDIA Blackwell adapter:
detBin=2/4/8corrected-frame checksums: exact against the zero-bad-before-bin referencefull-load low8 page profiles:
1.199/1.212/1.106 scrop-256 20-repeat medians:
0.774/0.755/0.733 scrop-256 p95:
0.798/0.813/0.775 snative non-low8
uint16detBin=2: exact at2.651 s
Local true 1024x1024x192x192 WebGPU product-first selected-block BF signoff on a real NVIDIA Blackwell adapter:
BF radius:
30selected compressed payload:
6.88 GBoutput image:
4.19 MB4-run median wall/profile/product:
4.92/4.85/1.56 smax/mean absolute error:
0/0against an independent Python referencethis is product-first BF evidence, not full-stack no-bin browse/load
Local true 1024x1024x192x192 WebGPU full-stack no-bin browser browse was driven with the shape-explicit local-H5 harness and rejected as a memory-policy path:
no crop, no detector bin, count-audited
uint8browse decode105 HDF5 data files plus 105 metadata-only block-index sidecars
selected corrected-frame CUDA checksums were prepared for first/middle/last frames
GPU memory reached about
97.2 GBof a97.9 GBbudgetbrowser failed before publishing a load profile/checksum readback with an invalid WebGPU buffer after a previous device/buffer error
do not mark strict full-stack browser browse as signed off for true 1024; use product-first, true scan crop, or explicit detector bin instead
Local real-data WebGPU browser stress on a 128x128 scan, 96x96 detector, uint8 sidecar, real NVIDIA Vulkan adapter:
mount/decode to interactive:
1.94 sBF warm recompute:
2.7-3.7 msDPC row warm recompute floor:
6.4-8.7 mswith browser scheduling outliersDPC col warm recompute floor:
8.3-9.7 mswith browser scheduling outliersidle RAF:
60.0 FPS
Exact focused parity tests for CUDA selected and dense-mask paths.
Local full 512x512x192x192 real-data benchmark:
BF median
4.96 ms -> 1.35 msADF median
16.16 ms -> 3.86 msDF median
62.64 ms -> 1.84 msmax absolute error
0for every row.
Local seven-tilt detector-bin2 real-data benchmark:
BF median
1.50 ms -> 0.54 msADF median
3.83 ms -> 1.35 msDF median
15.88 ms -> 0.53 msmax absolute error
0for every row.
Local seven-tilt Show4DSTEM compare-grid method benchmark after widget backend reuse:
BF full grid
3.97 ms(0.57 ms/panel)ADF full grid
9.56 ms(1.37 ms/panel)DF full grid
4.00 ms(0.57 ms/panel)max absolute error
0after the widget’s detector-area normalization.
Local DPC/CoM real-data benchmark:
full 512x512x192x192 uint16:
200.42 ms -> 12.39 ms, max absolute error0seven detector-bin2 panels: old summed median
373.14 ms; uncached CUDA summed median24.06 ms; first backend-filled grid24.63 ms; cached repeat path returns the stored CoM arrays without launching a GPU kernel.
Failed Hypotheses#
Replacing the full-detector CoM loop’s per-pixel detector-coordinate division with incremental row/column bookkeeping was exact but slower: the seven-panel DPC grid regressed from about
27.3 msto about34.5 ms. Keep the simpler division form unless a profiler points elsewhere.
Next Checklist#
Run Mac MPS
uint8runtime parity/performance on full no-bin data.Run Mac MPS CoM/DPC runtime parity/performance on full no-bin data and record the first-click vs cached-repeat timing.
Add a headed WebGPU test that records the real adapter and verifies BF/ADF/DF parity plus
maskedSumBufferdrag behavior.Keep WebGPU CoM/DPC/iDPC headed parity/performance tests in the release gate, including the FFT command-batching path that keeps iDPC median redraw under the 30 FPS budget.
Remove any remaining widget-local permanent backend copies after each synced domain-owned WebGPU source is covered by build and browser parity tests.
Treat real WebGPU
1024x1024x192x192full-stack no-bin browse/load as explicitly rejected unless a future browser/device path avoids materializing the whole decoded stack. Product-first BF for true 1024 is already signed off.