Memory management#
Use this page when you care about how much RAM/VRAM a dataset takes, which GPU runs the work, and how to give the memory back when you are done. For opening files and choosing a loader, start with IO/GPU.
Dtype in plain language#
The dtype is how each number is stored.
dtype |
range |
size |
use it for |
|---|---|---|---|
|
0 to 255 |
1 byte |
fast preview copies |
|
0 to 65535 |
2 bytes |
raw detector counts |
|
decimals |
4 bytes |
processed maps |
For raw electron detector counts, start with uint16. It keeps the measured
counts exactly and is still much smaller than float32.
Use uint8 only when you want a lightweight preview or tutorial copy. It is
fast and small, but it can saturate real counts above 255. load always keeps
the detector’s own dtype; uint8 appears only where you choose it, such as
export_html(dtype="uint8") or quantem show4dstem ... --html --dtype uint8.
Size estimates#
A 4096 x 4096 image is about:
dtype |
size |
|---|---|
|
16 MB |
|
32 MB |
|
64 MB |
A common 512 x 512 x 192 x 192 4D-STEM scan is 18 GiB as a dense uint16
array. load never creates that array: it keeps the acquisition ANS encoded on
the GPU at full detector resolution.
form |
approximate size |
|---|---|
dense |
18 GiB |
encoded acquisition from |
0.1 to 2 GiB, depending on counts |
bounded |
2.25 GiB |
Leave a few GB free for the viewer, browser, and downstream processing.
NVIDIA GPU workflow#
Most lab workflows should run Python on the NVIDIA workstation and open JupyterLab from a laptop. The workstation holds the data and runs the GPU work; the laptop is the frontend.
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
loaded = load("scan_master.h5") # CUDA is selected automatically when available
Show4DSTEM(loaded)
The same call fits every common GPU size, because the encoded acquisition is a small fraction of the dense array. Memory pressure comes from what you read or reconstruct from it, such as large bounded reads or SSB workspaces.
Check the GPU before and after a large load with quantem.widget.io.memory():
from quantem.gpu.io import load
from quantem.widget.io import memory
memory() # check VRAM before loading
loaded = load("scan_001_master.h5", verbose=True)
print(loaded.shape, loaded.dtype,
f"{loaded.resident_bytes / 2**30:.2f} GiB encoded, {loaded.logical_bytes / 2**30:.1f} GiB dense")
memory() # confirm VRAM after loading
For a real 256 x 256 x 192 x 192 Arina scan (uint32 counts on disk), the
print line reads:
(256, 256, 192, 192) uint16 0.57 GiB encoded, 4.5 GiB dense
Read this as: every detector pixel and count is on the NVIDIA GPU in encoded
form, 0.57 GiB instead of 4.5 GiB; the counts fit in uint16, so they are
stored that way, and no copy has been binned or quantized. The memory() lines
show the whole GPU, including other processes.
Can I choose the NVIDIA GPU inside the notebook?#
Yes. Put this in the first notebook cell, before importing torch, cupy,
quantem.widget, or any other GPU package:
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "0" # use physical NVIDIA GPU 0
Then import and load normally:
import torch
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
print(torch.cuda.get_device_name(0))
loaded = load("scan_001_master.h5")
Show4DSTEM(loaded)
Example output on a Linux workstation with NVIDIA GPUs:
cuda available: True
visible device count: 1
notebook device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
free 80.3 GiB / total 94.9 GiB
Inside that notebook, the selected GPU is called cuda:0. CUDA renumbers the
visible device, so physical GPU 1 also appears as cuda:0 if you selected it
with CUDA_VISIBLE_DEVICES="1".
How do I switch from GPU 0 to GPU 1?#
Change the first cell, restart the kernel, then run from the top:
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "1" # switch to physical NVIDIA GPU 1
Example output after restarting the Python process with GPU 1 selected:
cuda available: True
visible device count: 1
notebook device 0: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
free 94.4 GiB / total 95.0 GiB
Restarting matters. Once CUDA is initialized in a Python process, changing
CUDA_VISIBLE_DEVICES later in the notebook is not a reliable way to move the
work to another GPU.
How do I use two NVIDIA GPUs at the same time?#
Run one Jupyter process per GPU. Start each server with a different
CUDA_VISIBLE_DEVICES value:
# terminal 1: GPU 0
CUDA_VISIBLE_DEVICES=0 jupyter lab --no-browser --ip=0.0.0.0 --port=8888
# terminal 2: GPU 1
CUDA_VISIBLE_DEVICES=1 jupyter lab --no-browser --ip=0.0.0.0 --port=8889
Then open the printed URLs from your laptop. Each notebook sees its assigned
GPU as cuda:0.
Check what the notebook sees:
import torch
print(torch.cuda.is_available())
print(torch.cuda.get_device_name(0))
print(torch.cuda.mem_get_info()) # free bytes, total bytes
Check and free GPU memory#
Check the GPU before and after a large load:
import torch
free, total = torch.cuda.mem_get_info()
print(f"free {free / 1e9:.1f} GB / total {total / 1e9:.1f} GB")
Keep handles to the acquisition and the viewer if you plan to release memory later:
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
loaded = load("scan_001_master.h5")
viewer = Show4DSTEM(loaded)
viewer
When you are done with that dataset, release the viewer first and then the acquisition it borrows:
viewer.free() # releases widget tensor/backend caches
viewer.close() # closes the ipywidget comm/model
loaded.close() # returns the encoded storage to the GPU
A viewer from Show4DSTEM.from_folder(...) owns the acquisitions it loaded;
viewer.free() or viewer.close() closes them.
If memory is still occupied after this pattern, another variable, notebook, or kernel still owns it. A small residual allocation can remain because CUDA keeps a runtime context and small caches alive until the kernel exits. Shut down old kernels from JupyterLab before assuming the GPU is stuck.
Moving image data to Torch or CuPy#
Most viewers accept NumPy arrays or quantem datasets directly, so you usually do not need to move a PNG, TIFF, or EMD survey image to Torch just to view it. Move data to the GPU when you are about to run your own GPU computation.
For Torch:
import numpy as np
import torch
from quantem.widget import read_image
ds = read_image("haadf.emd")
image = np.ascontiguousarray(ds.array, dtype=np.float32)
device = "cuda" if torch.cuda.is_available() else "cpu"
image_t = torch.as_tensor(image, device=device)
For a large CPU array that you will reuse many times on an NVIDIA GPU:
if torch.cuda.is_available():
image_t = torch.from_numpy(image).pin_memory().to("cuda", non_blocking=True)
torch.cuda.synchronize()
For CuPy:
import cupy as cp
image_gpu = cp.asarray(image)
Keep raw detector counts as uint16 until you need decimal math. Convert to
float32 for filtering, fitting, normalization, neural networks, or display
processing. Avoid accidental float64; it doubles memory with no benefit for
normal interactive viewing.
For large .npy files, memory-map first so Python does not copy the whole file
before you decide what to view:
import numpy as np
stack = np.load("stack.npy", mmap_mode="r")
preview = np.asarray(stack[::8], dtype=np.float32) # explicit preview reduction
Apple Silicon workflow#
On a MacBook, the same API works:
from quantem.gpu.io import load
from quantem.widget import Show4DSTEM
loaded = load("scan_master.h5") # Apple GPU (MPS), encoded like on CUDA
Show4DSTEM(loaded)
Mac unified memory is shared by the operating system, browser, Python, and GPU. The encoded acquisition uses a small part of it; large bounded reads and reconstructions use the rest, so read the scan region you need rather than the whole scan.