Backend Optimization Matrix#
This page is the single checklist for accelerated Show4DSTEM and SSB work across CUDA, MPS, and WebGPU. It summarizes what is optimized today, what is only source-present, and what still needs real user-facing agreement and FPS signoff.
Do not count a speedup if it changes the microscope evidence: no hidden scan crop, detector binning, BF reduction, saved derived cache, or CPU fallback.
Frame Budgets#
Target |
Budget |
Use case |
|---|---|---|
Smooth drag |
|
60 FPS ROI dragging and scrub feedback. |
Interactive |
|
30 FPS BF/DF/ADF/DPC and SSB live steering. |
Reviewable |
|
10 FPS large or exact paths that are still usable. |
Slow path |
|
Needs redesign before calling it interactive. |
Report cold setup separately from warm interaction. For browser workflows, split file parse, decompression, upload, compute, readback, colormap, and canvas present. One end-to-end number is useful only after those stages are known.
Memory Footprint#
Raw 4D-STEM footprint for native 192x192 detector data:
Scan shape |
Raw |
Raw |
One float32 image |
|---|---|---|---|
|
|
|
|
|
|
|
|
SSB Hermitian G_qk footprint, complex64 half-plane:
Active BF |
|
|
|
|
|---|---|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Full-plane SSB G_qk is about twice the Hermitian footprint and is not a
public runtime mode.
Streaming Ptychography IO#
No-bin iterative ptychography should not require the full raw 4D/5D evidence to be resident in VRAM. Keep calibration and reconstruction stages streaming:
Compute COM, detector-center, rotation, and BF/DF/DPC calibration products as chunked reductions. Store the small maps or fitted values, then release the raw detector chunk.
Treat this first raw stream as cache generation. The interactive screen path should load the small BF/DF/CoM/rotation cache and must not recompute exact full-field products from raw HDF5 on every open.
Load the acquisition once with
io.load, which keeps it complete and ANS encoded, then feed the solver bounded rectangles withread(scan_region=...)from that resident.Auto-tune solver batch size to VRAM. A full
1024x1024x192x192 uint16acquisition is about77 GB; a1000x192x192 uint16batch is about74 MBbefore ptychography float/complex working buffers.
This is the path that makes strong no-bin ptychography plausible on 24 GB GPUs: VRAM holds the encoded acquisition, the current mini-batch, probe/object patches, propagators, gradients, and optimizer state, not the dense acquisition.
Product Coverage#
Product path |
CUDA |
MPS |
WebGPU |
Current gap |
|---|---|---|---|---|
HDF5 load/decompress |
CUDA bitshuffle/LZ4 kernels into one ANS-encoded resident, bounded |
Metal bitshuffle/LZ4 into one ANS-encoded resident in unified memory, bounded |
Browser HDF5/chunk reader, local-file acquisition, explicit detector-bin source path, selected-block sidecars, and WGSL decode sources. |
WebGPU full |
BF/ADF sparse masks |
RawKernel selected-pixel reducer with warp-shuffle reductions. |
Metal selected-pixel reducer on chunk-backed |
|
WebGPU needs broader repeated drag FPS signoff beyond the current full-size gates. |
Dense DF |
Cached full-detector total minus complement. |
Cached total minus complement. |
Cached full-detector total minus complement source path. |
WebGPU needs explicit dense-DF timing and memory behavior at |
CoM/DPC |
Fused CUDA moment reducer, backend cache. |
Raw Metal |
|
WebGPU DPC row/col clears the strict |
iDPC |
Uses shared DPC phase reconstruction after CoM. |
Same API over MPS CoM outputs. |
Fixed-rotation browser solver with paired DPC buffers and a dual-real FFT. |
WebGPU iDPC is signed off at float32 FFT tolerance and clears 30 FPS by median; p95/outlier tuning remains. |
SSB object redraw |
Optimized native kernels for |
Implemented; real |
ShowPtycho SSB WGSL source supports |
MPS/WebGPU need broader same-BF real-data signoff. |
SSB exact phase/loss |
|
|
Source exists; not yet a CUDA-level matrix. |
Extend 12-cell matrix with real WebGPU adapter and MPS full-BF runs. |
Current Measured Checkpoints#
MPS exact load sprint, 2026-08-22#
This sprint measured the earlier dense MPS loader. io.load now keeps the
native detector and returns an ANS-encoded resident, so the detector-bin rows
below are history.
The accepted path keeps the complete 512x512 scan, native uint16 counts,
scan bin 1, no crop, and explicit detector bins 1/2/4/8. All output bytes match
the pre-optimization path. Source pages were warm and uncontrolled; these are
library load measurements on MacBook Pro (M5 Max, 128 GB), not cold storage or
Live4DSTEM E2E.
Hypothesis |
Measured result |
Decision |
|---|---|---|
Decode full |
Bin1 ABBA p50 improved from 0.689 to 0.523 s; driver allocation sampled after load was 18.442 GiB for an 18.00 GiB logical resident. |
Promote. |
Fuse bit-unshuffle and exact detector summation for bins 2/4/8, with a specialized bin-2 kernel. |
Candidate p50 reached 0.498/0.421/0.417 s for bins 2/4/8; full-output hashes are byte exact. |
Promote. |
Allocate LZ4 scratch only for chunks that need it. |
Reduced persistent decoder allocation without changing output; retained as part of the accepted topology. |
Promote. |
Use 64 threads for the full exact decoder. |
Neutral to slower than 128 threads in the retained ABBA trial. |
Reject as default; retain artifact. |
Use 256 threads for the full exact decoder. |
Neutral to slower than 128 threads in the retained ABBA trial. |
Reject as default; retain artifact. |
Increase whole-output grouping beyond the accepted source-shard-aligned plan. |
A 3 GiB grouping regressed; 0.75 and 1.0 GiB were only marginal alternatives. |
Keep the 1.5 GiB compact-output default and the 1 GiB unusual-shard safety bound. |
Retune LZ4 threadgroup height to 12 or 16. |
No accepted end-to-end win over the retained layout. |
Reject as default; retain artifacts. |
One instrumented bin1 run measured 0.622 s wall, 0.211 s cumulative source reads, 0.390 s GPU interval union, and 0.093 s of gaps inside the GPU span. Source reads and GPU work overlap and must not be summed. The strict 0.5-second four-bin goal is met for bins 2/4/8; bin1 remains 23 ms above it by p50.
MPS compressed-save sprint, 2026-07-26#
Target workflow: full no-bin 512x512x192x192 MAPED output on Apple Silicon,
standard Arina master/data HDF5 layout, one frame per HDF5 chunk, Bitshuffle/LZ4
filter 32008, no detector binning, no scan crop, and no preview cache standing
in for raw detector evidence.
Public API rule: notebooks and user workflows should call quantem.gpu.io.save
with backend="auto", dtype=..., and the default Bitshuffle/LZ4 compression.
save_compressed_arina_h5 is the internal portable writer used by the public
CPU/MPS paths and directly exercised by backend maintenance tests. An Apple Silicon public-API validation using
from quantem.gpu import io; io.save(..., backend="auto", dtype="u16")
measured 1.91 s save, 3.14 s load+save, 1.205 GB output, and 512/512
exact decoded samples on the full no-bin MAPED master.
Committed path:
Commit |
Change |
Full-data result |
|---|---|---|
|
Added MPS/CUDA |
Full |
|
Overlapped MPS compression with HDF5 |
Full |
|
Tuned the MPS LZ4 hash table and added repeated-byte fast paths for bitshuffled data. |
Full |
|
Added a repeated-byte speed encoder for the exact-count MPS path. |
Exact |
|
Added native Metal chunk-backed |
Exact |
|
Re-tuned the native MPS save default batch from 4096 to 2048 after a full-data sweep. |
Exact |
|
Made |
Public API validation on the reference MPS host used the same full no-bin source and saved in |
Implementation checklist:
Item |
Status |
Evidence / next action |
|---|---|---|
Public load entry point |
Done |
|
Public save entry point |
Done |
|
One public compression method |
Done |
User-facing docs present Bitshuffle/LZ4 only; alternate codecs stay in internal archival helpers. |
MPS exact |
Done |
Native Metal chunk-backed path, full no-bin save |
MPS |
Done |
Full no-bin save |
MPS |
Done |
Implemented and covered by synthetic exact round-trip; slower than integer save and not the demo default. |
CUDA compressed save |
Done |
Existing CUDA writer remains the CUDA path for CuPy arrays. |
Portable reference save |
Done |
|
Public API collision guard |
Done |
Regression test keeps |
Same-real-MAPED CUDA-vs-MPS save timing |
Open |
Needed before publishing exact cross-backend save-speed claims. |
End-to-end seven-tilt stream/merge/save under 20 s |
Open |
Save stage is inside target; remaining work is alignment/merge pipeline overlap and single-pass acquisition staging. |
WebGPU HDF5 writing |
Deferred |
Browser writing is not part of the microscope MAPED path; keep WebGPU focused on review/product-first workflows for now. |
Final Apple Silicon full-run results:
Path |
Load |
Save |
Load + best save |
Output size |
Agreement gate |
|---|---|---|---|---|---|
|
|
|
|
|
4096 random frame/pixel samples exact in the earlier full gate; default-path 512 decoded samples also exact, mismatches |
|
|
|
|
|
4096 random frame/pixel samples exact versus explicit |
Backend gap/use-case summary:
Backend |
Current best role |
Save/write status |
Remaining gap |
|---|---|---|---|
CUDA |
Workstation/reference GPU path. |
|
Need a same-real-MAPED timing run to publish exact CUDA-vs-MPS save numbers. |
MPS |
Apple Silicon and microscope-local Mac workflow. |
Full no-bin |
Further wins are likely load/save pipeline overlap, lower file-size tuning, or moving more MAPED merge output directly into chunk-backed Metal buffers. |
WebGPU |
Browser review, local-file interaction, and front-end products. |
HDF5 writing is intentionally a gap. |
Useful later for browser-only export/share without Python; not needed for the current microscope MAPED processing path. |
Bounded experiments from this sprint:
Hypothesis |
Result |
Decision |
|---|---|---|
Feed |
Correct, but full |
Do not use direct-MLX staging as the default. |
Native Metal compressor with per-batch scratch allocation. |
Correct, but compress+pack took about |
Reuse scratch buffers instead. |
Native Metal compressor with reusable scratch. |
Compress+pack for all full |
Promote for chunk-backed MPS |
Native Metal batch-size sweep. |
Full-data saves measured 2048 frames/batch at |
Use 2048 as the MPS save default. |
RLE-only LZ4 encoder for exact |
Faster than the hash encoder, but output grew from about |
Accept for the speed path because output remains portable standard LZ4 and keeps the demo under 2 s. |
Workflow |
Backend |
Shape / BF policy |
Result |
Status |
|---|---|---|---|---|
BF virtual image |
CUDA |
full |
|
Strong. |
ADF virtual image |
CUDA |
full |
|
Strong. |
Dense DF virtual image |
CUDA |
full |
|
Strong. |
CoM/DPC |
CUDA |
full |
|
Strong. |
HDF5 load/decompress |
CUDA |
full |
cold first load about |
Meets single-dataset warm target; cold includes process/device/cache startup. |
HDF5 load/decompress |
CUDA |
true real |
|
Strong CUDA reference signoff for true 1024 acquisition. |
Stochastic HDF5 ptycho minibatch |
CUDA, reference GPU |
40 real-data masters, 1000 global random scan positions per master, detector |
historical internal worker sweep at |
Correct global stochastic order and raw counts. The public API exposes no worker knob and uses an internal bounded single-reader scheduler. Bottleneck is scattered HDF5 payload access, not bitshuffle/LZ4 GPU decompression. |
BF/DF/CoM/rotation cache build |
CUDA, reference GPU under 12 GB allocator cap |
true real |
first build |
Demonstrates no-bin full-field calibration products on a 12 GB-style budget. This is the product-cache build path. It is too slow for page launch and must not be presented as the interactive path. |
BF/DF/CoM/rotation cache build |
MPS on Apple Metal |
true real |
cache build |
Mean DP/BF/DF bit-exact versus CUDA; CoM row/col max abs error |
BF/DF/CoM/rotation cache hit |
Any backend-facing caller |
true real |
local cache read repeats |
This is the screen/UI launch path for BF/DF/DPC/rotation. Cache hits are backend-neutral and do not probe CUDA before returning. |
HDF5 load/decompress |
MPS |
true real |
|
Strong MPS reference signoff for true 1024 acquisition; close to the conservative Apple memory guard and should stay chunk-backed. |
Seven-master HDF5 load/decompress |
CUDA |
seven full |
explicit warmup then measured loads sum to |
Meets |
Seven-panel BF/ADF/DF grid |
CUDA |
detector-bin2 seven-panel real workflow |
BF |
Strong for current grid policy. |
MPS no-bin load + VI + CoM smoke |
MPS |
full |
load about |
Smoke passed; formal crop-product agreement now covers the masked CoM fallback. |
HDF5 load/decompress |
MPS |
full |
the public load at that revision returned |
Near CUDA steady-state for one load; retained seven-load target is close, independent cold-ish repeats still show Apple memory-pressure variance. |
Show4DSTEM WebGPU headed stress |
WebGPU |
|
mount/decode about |
Promising small case; not a full |
HDF5 load/decompress |
WebGPU |
full |
best URL total |
Functional full- |
HDF5 load/decompress |
WebGPU on Apple GPU |
full |
full-stack total |
Hardware path is valid, but tunneled HTTP acquisition dominates; use local-file acquisition for real local review. |
HDF5 local-file Show4DSTEM production path |
WebGPU on Apple GPU |
full |
previous general |
Strong browser checksum parity versus CUDA and about a |
WebGPU detector-bin local-file load |
WebGPU on NVIDIA Blackwell |
full |
headed Chrome low8 browse page profiles |
Strong explicit detector-bin signoff, including crop+bin repeated coverage. A widget regression was fixed so raw detector bad-pixel indices are not reapplied to binned output pixels after the load path has already zeroed raw bad pixels before binning. |
HDF5 local-file scan-region full-stack path |
WebGPU on Apple GPU |
true |
crop-aware frame-window decode with data-file prefilter and optional block-index sidecars: 946-cycle soak median |
Strong crop-first full-stack parity and now faster than the CUDA warm crop reference previously measured around |
Product-first BF from HDF5 |
WebGPU on Apple GPU |
true |
3-repeat median wall |
Strong crop-region parity. This is not a prefix; row-major scan-region mapping is explicit. |
Product-first BF from HDF5 |
WebGPU on Apple GPU |
full |
3-repeat median wall |
Strong product-first parity without materializing the |
Product-first BF selected-block sidecar |
WebGPU on Apple GPU |
true |
auto pixel/direct-float/staging-pipeline kernel: 946-cycle soak median |
Strong crop parity. The crop path uses sidecar span filtering before full file reads and auto-batches small row-window specs. |
Product-first BF selected-block sidecar |
WebGPU on Apple GPU |
full |
auto grouped-mask/direct-float/staging-pipeline kernel: 946-cycle soak median |
Current best. Hits the CUDA-like single-product target for this BF product by storing exact selected detector-block streams instead of reading unrelated HDF5 detector blocks. |
Product-first BF selected-block sidecar |
WebGPU on NVIDIA Blackwell |
true real-acquisition |
4-run median wall |
Strong product-first true-1024 parity on a real WebGPU adapter. This does not claim full-stack browser browse/load signoff because it intentionally reads only BF-touched detector blocks and does not materialize the |
Product-first BF selected-block sidecar |
WebGPU on Apple GPU |
|
auto grouped-mask/direct-float/staging-pipeline kernel with compact shared memory for large scans: 946-cycle soak median |
Dispatch/output scaling gate only. Not a true 1024 real-acquisition signoff. |
WebGPU load/product soak |
WebGPU on Apple GPU plus CUDA reference on RTX PRO 6000 Blackwell |
full |
946 cycles produced 5676 timing rows. Five rows had transient Chrome/CDP socket or timeout harness failures; successful parity rows had no numeric mismatch. |
Use this as the current stability baseline for WebGPU IO/product changes. It is not a substitute for true 1024 acquisition signoff. |
Product-first BF from HDF5 |
WebGPU on Apple GPU |
|
3-repeat median wall |
Dispatch/output scaling gate only. Not a true 1024 real-acquisition signoff. |
HDF5 load/decompress experiment |
WebGPU |
full |
sidecar generation |
Not adopted as a default speed path; high |
BF virtual image |
WebGPU |
full |
first BF r30 click about |
Warm path is interactive; first click still includes setup/cache work. |
Dense DF virtual image |
WebGPU |
full |
first click about |
Uses total-minus-complement path, but complement reducer still needs CUDA-level tuning. |
DPC/CoM/iDPC |
WebGPU |
full |
corrected-frame load parity passed; DPC row/col/iDPC display medians |
Strong full no-bin DPC row/col signoff. iDPC now clears 30 FPS by median after batching FFT command submissions; p95/outlier and float32-FFT-tolerance tightening remain. Benchmark artifacts must reject URL fallback for local-file timing claims. |
Real HDF5 crop-first equality gate |
CUDA |
full |
old/new full checksum passed; crop-first data exactly matches full-load slice |
Strong IO/decompress parity gate on real data. |
Real HDF5 crop-first equality gate |
MPS |
full |
crop-first data exactly matches full chunked slice on Apple GPU |
Strong MPS IO/decompress parity gate on real data. |
Real crop product agreement gate |
CUDA |
opt-in local HDF5 crop, BF radius |
BF/ADF/DF exact; full and BF-masked CoM within |
Strong product correctness gate; wall time is reference-heavy, not a clean benchmark. |
Real crop product agreement gate |
MPS |
opt-in local HDF5 crop, BF radius |
BF/ADF/DF exact; full and BF-masked CoM within |
Strong product correctness gate on Apple GPU. |
WebGPU corrected-frame checksum gate |
WebGPU on Apple GPU |
full |
selected-frame |
Strong browser HDF5 parse/decode/chunk-order/dtype/bad-pixel parity gate without reading the full 9.7 GB stack back to CPU. |
SSB exact phase/loss |
CUDA |
real |
mean about |
Meets 30 FPS with small p95 margin. |
SSB exact phase/loss |
CUDA |
synthetic |
about |
Not interactive yet. |
SSB exact phase/loss |
MPS |
real |
fresh Apple Silicon |
Correct and reviewable at about |
SSB exact phase/loss |
MPS |
real |
fresh Apple Silicon |
Correct but slow. This is the large-BF policy that still needs a deeper MPS row/column FFT topology. |
SSB exact phase/loss |
MPS |
synthetic full-BF-style |
fresh Apple Silicon source-tree probe: object median |
Needs deeper MPS topology for large exact phase/loss; chunk tuning alone is not enough. |
Required Agreement Gates#
Every new optimization should add or update one of these gates before claiming speed:
Gate |
Reference |
Required metrics |
|---|---|---|
CUDA BF/ADF/DF |
Previous CuPy or NumPy selected-pixel sum on the same resident data. |
max abs error, dtype, mask pixel count, timing before/after, peak temp memory. |
CUDA CoM/DPC |
Previous CoM implementation on the same mask and scan. |
row/col max abs error, centered DPC error, timing before/after, cache behavior. |
MPS BF/ADF/DF |
NumPy or CUDA reference from the same loaded evidence. |
max abs error, first-click timing, warm repeated timing, unified-memory footprint. |
MPS CoM/DPC |
CUDA or NumPy reference with the same detector mask and coordinate convention. |
row/col max abs error, DPC component error, first-click timing, cached repeat timing. |
Real HDF5 crop products |
Independent NumPy reference from the exact loaded crop. |
BF/ADF/DF bit-exact raw-count sums; full and masked CoM error at |
WebGPU BF/ADF/DF |
Python |
browser adapter, max/mean/p99 error, warm compute ms, display ms, no SwiftShader timing claims. |
WebGPU CoM/DPC/iDPC |
Python CUDA/MPS reference arrays with same mask, mean policy, and fixed iDPC rotation. |
row/col error, centered component error, iDPC float32 FFT error, compute/readback/display split. |
WebGPU HDF5 load/decode |
Python CUDA/MPS corrected-load reference from the same local evidence. |
selected corrected frame checksum, frame order, dtype, bad-pixel count, adapter, load split, no private path in artifacts. |
SSB object |
Corrected-object reference at same BF, aberrations, scan size, and precision. |
object complex error or phase/amplitude image error, warm redraw FPS. |
SSB phase/loss |
Existing exact phase/loss path; never object-wave identity as reference. |
phase mean/p99/max, scalar loss delta, fit/reconstruction agreement. |
Next Optimization Queue#
WebGPU product-first HDF5 evidence layout: exact BF r30 products now pass CUDA parity for true
256, full512,1024repeat-stress, and true real-acquisition1024product-first evidence. The selected-block sidecar path proves the0.5 sfull-512target is reachable when the browser reads only exact detector-block evidence. The production local-H5 masked-sum API now discovers sidecars, checks selected-block coverage, filters crop spans before full reads, and falls back honestly when a dragged ROI needs detector blocks not present in the sidecar. Next production work is maintaining the sidecar cache writer/invalidation policy and wiring the fast product path into the Show4DSTEM UI product loop.WebGPU compressed-byte upload and decoder floor: the count-audited low8 frame-cooperative kernel plus staging pipeline cut the full local-file path to a
0.725 spage profile in the latest full-512block-index run. The selected-block full-512BF product path reaches0.370 spage total /0.209 sproduct stage with exact parity in the latest fresh Chrome run. The remaining full-stack gap is materializing the whole9.7 GBbrowse cube; the product path shows why selected evidence is the right interactive route.WebGPU pipelined selected-block acquisition: extend the staging-pipeline idea to browser range/local cache management, with a bounded staging/raw buffer ring and explicit peak-memory reporting.
WebGPU Show4DSTEM full-
512report: load/decode/upload/compute/display split for BF, ADF, dense DF, CoMx, CoMy, CoM magnitude, and browser iDPC. DPC row/col has a full no-bin headed signoff; iDPC has a float32 FFT tolerance signoff and a median 30 FPS pass after FFT command batching. Continue tightening p95/outliers.WebGPU bitshuffle/LZ4 kernel redesign: the frame-cooperative low8 path is now the default for count-audited lossless
uint8browse loads. The next jump needs lower atomic/shared-memory pressure or a different browser payload contract, not more fetch batching.WebGPU full-
512agreement harness: compare browser outputs againstquantem.gpuCUDA/MPS references with public-safe fixture labels and image difference maps.WebGPU generated-HDF5 browser gate: cover
h5_urlfetch, HDF5 parse, bitshuffle/LZ4 decode, bad-pixel masks, BF/DF/ADF, and DPC without relying on local fixture paths.WebGPU dense-DF cache signoff: verify
total - complementbehavior on a headed real adapter so large annular/dense masks do not scan most detector pixels on every drag.WebGPU timing telemetry follow-up: source now reports adapter info, resident bytes, fetch/parse/upload/build/GPU-wait splits, and timestamp-query availability; next step is optional GPU timestamp pass timing.
WebGPU iDPC display path: fixed-rotation Poisson integration is implemented with paired DPC buffers and a dual-real FFT. Continue optimizing toward full no-bin 30 FPS and keep the float32 FFT parity gate in the browser benchmark.
MPS full-
512VI/DPC report: BF/ADF/DF/CoM/DPC first-click and warm-repeat timings with the same masks as CUDA references.MPS IO signoff: encoded no-bin load, uint32 narrowing, fast sidecar decode, bounded
readrectangles, and expected memory-guard failures.WebGPU detector-bin breadth: full-512 and true crop-256
detBin=2/4/8are signed off with exact corrected-frame checksums and repeated/p95 crop evidence. Keep detector-bin presets explicit and refresh p95 after any decoder/upload changes.WebGPU true
1024x1024x192x192acquisition: CUDA and MPS real full-stack loads are signed off, and WebGPU product-first BF selected-block signoff is exact on true real-acquisition evidence. Browser full-stack no-bin browse/load still needs either enough free WebGPU VRAM for a true full-stack run or an explicit documented memory-policy rejection.SSB MPS
512/1024exact phase/loss topology: optimize row/column FFT work; chunk-size-only changes are not enough. Fresh Apple Siliconorigin/mainprobes show radius-30512exact phase/loss around76 ms, full-active-BF512around529 ms, and full-BF-sized synthetic1024around669 ms; the usable object-wave steering path is a separate quantity.SSB WebGPU matrix: run
128/256/512/1024object, phase, and phase+loss against CUDA references on a real WebGPU adapter.
Failed / Bounded Hypotheses#
Hypothesis |
Result |
Decision |
|---|---|---|
Increase WebGPU HDF5 decode batch from |
Kernel/decode stage can improve ( |
Keep default batch |
Increase data-file fetch window from |
Higher windows reduced some fetch waits in isolated runs but caused master/data contention or decode variance; total stayed worse than default. |
Keep default fetch window |
Fetch/parse HDF5 master before data files. |
Master timing dropped to about |
Do not reorder as default. |
Embed HDF5 bad-pixel metadata and skip browser master fetch. |
Removes |
Keep for exported local H5-source fixtures; it is simpler and not a hidden evidence change. |
Production local-file worker/group/decodeBatch sweep. |
Real Chrome |
Keep default workers/group/batch |
Module Blob worker for local HDF5 reads. |
|
Use a classic inline Blob worker for local-file acquisition. |
Fused |
Valid browser path, but full-source compressed bytes remain about |
Keep fused |
Low8-only bslz4 decode. |
Count audit after bad-pixel correction showed max unmasked count |
Promote frame-cooperative low8 for count-audited |
HDF5 frame-index manifest only. |
The metadata-only frame-offset manifest preserved parity and reduced parse time, but full-stack local load improved only about |
Keep the manifest as a useful metadata cache and as a prerequisite for future selected-block/range layouts; do not claim it is the load-time fix by itself. |
Product-first masked-sum HDF5 decode. |
Exact BF r30 parity now passes for a true |
Promote as the correct virtual-image product direction. The next speed jump requires selected-block acquisition/layout, not more scan cropping or BF reduction. |
Selected-block sidecar layout. |
A sidecar preserving exact native bslz4 streams for selected detector blocks cut full |
Promote as the browser fast-load contract for product-first VI/DPC/SSB evidence, provided metadata records the detector-block coverage and the UI falls back when the mask asks for missing blocks. |
Sidecar crop prefilter plus auto product batching. |
Moving sidecar span filtering before full reads and auto-batching small row-window specs cut the true |
Promote for scan-region products. Keep full |
Pixel reducer workgroup variants. |
Forced |
Keep wider pixel variants as profiling knobs only. The production default is now the grouped-mask |
Grouped-mask popcount reducer. |
Re-tested after the current browser/source sync, grouped-mask |
Use automatic routing: pixel |
Direct-float selected-block output. |
Writing float32 product pixels directly from the masked-sum decoder removed the separate u32 output buffer and conversion pass while preserving exact parity. It improved full |
Keep direct-float output as the production selected-block product path. The sums are below float32’s exact integer range for the tested BF/DF masks. |
Selected-block staging pipeline. |
Preparing group |
Promote staging pipeline by default for selected-block product-first kernels. Leave |
Selected-block staging lifecycle guard. |
A true-1024 browser harness exposed illegal remapping of staging buffers before submitted copy work had completed ( |
Keep the staging-ready guard in production source. Product benchmarks should fail fast on missing fixture files, mounted-file count mismatch, URL fallback, and absent reference arrays. |
Large selected-block compact shared memory. |
Shrinking grouped-mask shared-memory clearing from full |
Promote only for scans above |
Block-parallel selected-block grouped-mask prototype. |
Splitting each selected detector block into separate workgroups and accumulating with u32 atomics preserved exact parity, but slowed the |
Removed from production source. The extra atomics/conversion pass did not pay for three selected BF blocks. |
Serial grouped-mask selected-block decoder. |
One-lane LZ4 decode followed by the same grouped-mask popcount preserved exact parity, but slowed full |
Removed from production source. The cooperative LZ4 copy still beats serial decode despite atomics and token-loop barriers. |
Scratch selected-group sidecar. |
A mask-specific sidecar storing only the exact low-bitplane byte groups reduced the full |
Do not promote as a package format yet. Revisit only with a reusable upload pool, browser persistent GPU cache, or a broader sidecar policy that still supports interactive mask changes. |
Full-stack WebGPU frame-coop workgroup size. |
After fixing the template so |
Promote |
Full-stack WebGPU file-group scheduler. |
With |
Promote size-aware low8 local-file grouping: |
Full-stack WebGPU decode batch. |
At file group |
Keep full-stack |
Full-stack WebGPU direct frame-index parser. |
Replacing per-frame |
Keep. It removes allocation churn but does not change the main upload/decode floor. |
Full-stack frame-index worker parse. |
Moving frame-index bslz4 metadata construction into the local-file worker preserved checksum parity and, with full-load worker count |
Promote: use worker count |
Full-stack HDF5 block-index sidecar. |
A |
Promote as an optional metadata cache because it removes repeated HDF5 block-header walking. Do not claim it solves the strict |
Full-stack compressed sidecar with low-plane decode. |
A compressed-payload sidecar that stored normal low8 streams passed exact CUDA checksum parity but was not faster ( |
Removed from production source. The native bslz4 stream remains effectively upload-bound, and trimming decoded bit planes did not reduce compressed bytes enough to pay for another cache format. |
Full-stack WebGPU subgroup token parser. |
A scratch V-R shader removed redundant token parsing with subgroup broadcasts and looked fast ( |
Removed from production source. Do not promote any subgroup-token parser until it passes the corrected-frame checksum gate. |
Full-stack WebGPU worker count 4 refresh. |
Older retained artifacts had |
Keep the current default worker policy; do not flip back to 4 without a new paired win. |
Full-stack WebGPU branchless bit-transpose expression. |
Replacing the low8 bit-transpose loop with direct bit expressions preserved parity but slowed the default |
Reverted; the browser compiler handles the compact loop better on Apple WebGPU. |
Full-stack WebGPU combined staging upload. |
A single aligned staging/raw buffer per group preserved parity, but did not beat normal staging at the current default ( |
Do not promote. Keep normal per-spec staging as the default. |
Full-stack WebGPU packed-word shared low8 decoder. |
Replacing packed shared-memory byte atomics with owner-writes-full-word shared memory preserved corrected-frame checksum parity, but slowed full |
Do not promote. Removing atomics alone is not enough; the extra word assembly and shared-memory traffic outweighed the savings on Apple WebGPU. |
Full-stack block-cooperative low8 decoder template. |
The off-default block-cooperative template had an unreplaced shared-memory size placeholder that produced all-zero output in a control run. After fixing |
Keep the template fix for correctness of the profiling knob, but keep frame-cooperative |
Skip zero-literal token barriers. |
About |
Removed from production source; fewer barriers did not beat the compiler/GPU behavior of the uniform barrier path. |
Pack decoded output chunks into larger group buffers. |
Binding decoded chunks as offsets into about four larger output buffers preserved checksum parity but slowed the block-index full run to about |
Removed from production source; fewer output buffers did not offset allocation/zeroing/binding costs. |
Full-stack scan-region frame-window crop. |
Adding crop-aware bslz4 frame-window slicing and frame-index data-file prefilter made a true |
Promote for lower-level WebGPU local-H5 crop loads. UI wiring needs a matching scan-shape state path before using cropped stacks in Show4DSTEM display. |
Merge cropped frame windows into fewer full-stack specs. |
Merging the |
Removed from production source. The faster default keeps row-window specs and relies on decode batching/staging pipeline. |
Truncate native bslz4 uploads to the low8 prefix. |
Full token audit found the low8 prefix is |
Do not implement prefix-trim packing for native HDF5; it adds CPU token parsing with almost no byte savings. |
|
Full-load attempt did not finish in the normal profiling window and was not competitive with staging-buffer upload. |
Keep the staging uploader as the default and leave writeBuffer as an off-by-default experiment only. |
Chunked |
Splitting writes into |
Keep staging upload as default. Chunked writeBuffer remains a profiling switch only. |
Single-lane token parser for low8 WGSL decode. |
The first non-uniform-barrier version looked fast but produced all-zero frames, so it was invalid. A fixed |
Do not promote. A future parser must preserve uniform barriers without adding a per-token metadata barrier. |
Multiple frames per WebGPU workgroup. |
|
Keep default |
Aligned 32-bit word copies inside the WGSL LZ4 decode loop. |
Browser validated and rendered, but repeat full runs were not a stable speedup and sometimes slower. |
Reverted; next kernel attempt should be cooperative LZ4 copy, not extra scalar branches. |
Experimental parallel LZ4 WGSL kernel. |
Full |
Leave |
Direct mapped WebGPU compressed-byte upload. |
CPU upload bucket dropped ( |
Removed as production code; keep the staging uploader. |
Product-first WebGPU radial-profile HDF5 decoder for BF/DF/ADF. |
On real Chrome |
Removed from production source; pursue local-worker acquisition, pipelined decode, and cooperative LZ4 instead. |
MPS depth/read-ahead/LZ4-y/hazard sweeps. |
Default depth-2/read-ahead remained best or tied: retained-scratch medians stayed around |
Keep the simpler default MPS scheduler and scratch reuse. |