Performance
This page gives the time and memory OpenReadout needs, compared with Bio-Formats and bioio on the same files. The commands to reproduce every table are at the end. The raw results of the main comparison are in bench/results/compare.json.
Read the numbers with their conditions. They come from one laptop that was shared with other build jobs, so treat them as order-of-magnitude comparisons, not precise ratios. A † marks a measurement taken at a 1-minute load average above 3.
Conditions
| Machine | Apple M1 Pro (8 cores: 6 performance, 2 efficiency), 16 GB RAM, internal SSD |
| OS | macOS 26.2 |
| OpenReadout | 0.1.0 at commit ba314f2, cargo build --release, rustc 1.98.1 |
| Bio-Formats | bftools 8.5.0 (showinf, bfconvert), OpenJDK 17.0.5, default -Xmx512m |
| Python | CPython 3.12.12; bioio 3.5.1 with bioio-czi 3.0.0, bioio-nd2 2.5.0, bioio-lif 1.5.0, bioio-ome-tiff 1.4.0; flowio 1.4.0; pyabf 2.3.8; pyteomics 5.0.1 |
| Load | 1-minute load average 2.0–3.0 before most measurements |
Method
- Same files, warm cache. Each command runs once to warm the page cache, then 5 more times. The tables give the median. The files are from the test corpus.
- One process per run. Wall time is measured from the parent process. Peak resident memory (RSS) comes from
/usr/bin/time -l. - Start-up is measured separately. The Python readers also report the time spent inside the process after the imports (“in-process”), so interpreter start-up is visible but not counted twice.
- Like-for-like metadata.
openreadout info --json;showinf -nopix -no-upgrade; for bioio,BioImage(path)followed by the dimensions, data type, channel names, pixel sizes and metadata of every scene. - Like-for-like export. Uncompressed OME-TIFF from all three tools. OpenReadout’s time includes reading the whole output back and comparing every plane, which the others do not do.
Start-up (no file)
| Process | Time | Peak RSS |
|---|---|---|
openreadout --version |
5.2 ms | 4 MiB |
Bio-Formats showinf -version (JVM) |
221 ms | 64 MiB |
python -c pass |
18 ms | 15 MiB |
Python with import bioio and four plugins |
1.19 s | 174 MiB |
Metadata (info)
Wall time per process and peak RSS.
| File | Size | OpenReadout info |
Bio-Formats showinf -nopix |
Python reader (in-process time) |
|---|---|---|---|---|
CZI aics-s-3-t-1-c-3-z-5 |
15 MB | 7.7 ms; 7 MiB | 587 ms; 135 MiB | bioio 1.52 s (243 ms); 206 MiB |
CZI zenodo10577621-Young-mouse |
3.7 GB | 19 ms; 9 MiB | 991 ms; 322 MiB | bioio 8.22 s (6.90 s); 414 MiB |
LIF ome-imagesc-110520-AMR1 |
3.0 GB | 24 ms; 32 MiB | 945 ms; 288 MiB | bioio 2.47 s (1.26 s); 334 MiB |
ND2 ome-jonas-control002 |
48 MB | 6.4 ms; 5 MiB | 565 ms; 168 MiB | bioio 1.62 s (376 ms); 207 MiB |
OME-TIFF aics-s-3-t-1-c-3-z-5 |
9.7 MB | 11 ms; 14 MiB | 876 ms; 263 MiB | bioio 1.70 s (394 ms); 226 MiB |
FCS flowio-m0-wm278-s1-zero-gain |
13 MB | 6.5 ms; 5 MiB | n/a | flowio 73 ms (0.3 ms); 33 MiB |
ABF pyabf-2020-07-29-0062 |
0.8 MB | 4.9 ms; 5 MiB | n/a | pyabf 78 ms (2.3 ms); 35 MiB |
Thermo RAW mtbls404-QC1_001 |
42 MB | 7.9 ms; 10 MiB | n/a | n/a |
mzML mtbls404-QC1_001 |
115 MB | 50 ms; 6 MiB | n/a | pyteomics 1.93 s (1.45 s); 132 MiB |
For a one-off question about a file, a native binary answers in milliseconds because there is no runtime to start. JVM and Python tools spend most of their time starting up. Inside a running Python process, the FCS and ABF readers are as fast as OpenReadout. info on mzML reads each spectrum header through the index, so its time grows with the number of spectra.
Export to OME-TIFF
Wall time per process and peak RSS, with the default number of threads (8 here).
| File | OpenReadout (none) | OpenReadout (deflate) | bfconvert |
bioio save (in-process) |
|---|---|---|---|---|
CZI aics-s-3-t-1-c-3-z-5 (15 MB) |
18 ms; 22 MiB | 67 ms; 24 MiB | 2.56 s; 381 MiB | 1.57 s (425 ms); 232 MiB |
ND2 ome-jonas-control002 (48 MB) |
32 ms; 13 MiB | 237 ms; 18 MiB | 3.76 s; 286 MiB | 2.43 s (1.03 s); 359 MiB † |
LIF zenodo14976703-Convalaria-LambdaScan (15 MB) |
16 ms; 19 MiB † | 89 ms; 23 MiB | 1.35 s; 290 MiB | 1.54 s (350 ms); 221 MiB |
OME-TIFF aics-s-3-t-1-c-3-z-5 (9.7 MB) |
60 ms; 60 MiB | 115 ms; 62 MiB | 1.61 s; 306 MiB | 2.17 s (878 ms); 279 MiB † |
Threads
openreadout --threads N export FILE --compression …: median wall time, speed-up over 1 thread, and peak RSS. Planes are decoded and compressed on N threads and written by one writer. The output is byte-identical for every N.
| File | 1 thread | 2 threads | 4 threads | 8 threads |
|---|---|---|---|---|
ND2 ome-karl-sample-image 218 MB, none |
443 ms; 24 MiB | 229 ms (1.9x); 43 MiB | 157 ms (2.8x); 79 MiB | 137 ms (3.2x); 147 MiB |
| same, deflate | 4.66 s; 30 MiB | 2.46 s (1.9x); 48 MiB | 1.31 s (3.6x); 90 MiB | 1.06 s (4.4x); 157 MiB |
CZI openslide-zeiss-5-jxr 69 MB, none |
4.90 s; 409 MiB | 2.60 s (1.9x); 832 MiB | 1.66 s (3.0x); 912 MiB | 1.36 s (3.6x); 1054 MiB |
Deflate is limited by compression and gains the most from threads. Uncompressed export is limited by memory copies, the write and the read-back. The 8-thread column also uses the slower efficiency cores.
Memory
Export holds a bounded number of decoded planes, whatever the size of the file. Planes are decoded in windows of at most threads planes and at most 1 GiB, and at most two windows are alive at a time. So peak memory is a few planes plus a few tens of MiB, however large the file. Planes larger than 4 GiB are read and exported block by block instead of being assembled in memory.
crates/openreadout-bench/tests/memory_ceiling.rs enforces this. It counts heap bytes while exporting synthetic images of 16 and 64 planes at 1 and 4 threads, and checks that the peak does not grow with the number of planes.
Large images and whole slides
Release build, warm page cache, one run each, on the same laptop under heavy load (load average about 10).
Reading one 512 × 512 window at the centre of a level (stats --level L --region X,Y,512,512):
| File | Level 0 size | Level 0 window | Level 2 window |
|---|---|---|---|
zenodo10577621-Young-mouse.czi (3.7 GB, JPEG XR) |
190,309 × 69,378 RGB | 0.08 s, 59 MB | 0.05 s, 28 MB |
openslide-aperio-CMU-1.svs |
46,000 × 32,914 RGB | 0.01 s, 20 MB | 0.01 s, 20 MB |
openslide-hamamatsu-CMU-1.ndpi |
51,200 × 38,144 RGB | 0.02 s, 23 MB | 0.01 s, 20 MB |
ome-qptiff-HandEcompressed_Scan1.qptiff |
30,720 × 26,640 RGB | 0.01 s, 24 MB | 0.01 s, 18 MB |
Whole-slide exports; every block of every level is read back and compared before the rename:
| Export | Output | Time | Peak RSS |
|---|---|---|---|
| CMU-1.svs to OME-TIFF, source pyramid (3 levels) | 935 MB | 14.6 s | 1.04 GB |
| CMU-1.svs to OME-Zarr, source pyramid | 919 MB | 42.4 s | 1.20 GB |
| Young-mouse CZI, 20,000 × 20,000 region to OME-TIFF, 6 levels | 1.27 GB | 14.9 s | 1.15 GB |
Screening plates
info on a plate reads the index and lists the folder once. check also reads the TIFF header of every plane file. Release build, median of 3 runs, machine under shared load.
| Plate | info |
check |
|---|---|---|
| JUMP Phenix index (47 MB, 27,648 planes) | 335 ms; 101 MiB | 280 ms; 46 MiB |
| CellVoyager CV8000 (2,304 fields) | 107 ms; 111 MiB | 50 ms; 34 MiB |
| CellVoyager CV7000 (2,817 fields) | 72 ms; 63 MiB | 37 ms; 22 MiB |
Python arrays
bench/python_arrays.py runs each task in its own Python process, so the time includes importing the package and opening the file. Median of 3. full reads every plane of image 0 into one NumPy array; region reads the 512 × 512 window at the centre of full resolution. Both readers returned the same pixels in every row.
| File | Task | openreadout | bioio reader |
|---|---|---|---|
aics-s-3-t-1-c-3-z-5.czi |
full | 0.009 s, 38 MB | bioio-czi 0.968 s, 182 MB |
zenodo10577621-Kidney-RAC-3color.czi |
full | 1.26 s, 652 MB | bioio-czi 5.11 s, 1296 MB |
zenodo10577621-Young-mouse.czi |
region | 0.071 s, 78 MB | bioio-czi 8.42 s, 453 MB |
ome-karl-sample-image.nd2 |
full | 0.424 s, 256 MB | bioio-nd2 1.34 s, 612 MB |
aics-s-3-t-1-c-3-z-5.ome.tiff |
full | 0.036 s, 46 MB | bioio-tifffile 1.04 s, 203 MB |
Much of each small-file time is start-up. For regions, the gap is mostly because bioio’s plugins read whole planes, while OpenReadout reads only the tiles it needs.
Codecs
Single-threaded decode throughput on the same 64 chunks of 512 × 512 pixels from a brightfield slide, compared with imagecodecs, which wraps the C libraries. One run each, on a loaded machine.
| Codec | OpenReadout (crate) | imagecodecs (library) | Ratio |
|---|---|---|---|
| WebP lossy (RGB) | 300 MB/s (image-webp) |
564 MB/s (libwebp) | 0.53 |
| JPEG XL lossy (RGB) | 53 MB/s (jxl-oxide) |
179 MB/s (libjxl) | 0.30 |
| JPEG 2000 lossy (RGB) | 38 MB/s (rust-j2k) |
138 MB/s (OpenJPEG) | 0.27 |
| JPEG 2000 lossless, 16-bit grey | 6.5 MB/s | 21 MB/s | 0.31 |
| JPEG XR lossy (RGB) | 185 MB/s (openreadout-jpegxr) |
280 MB/s (jxrlib) | 0.66 |
| JPEG XR lossless, 16-bit grey | 120 MB/s | 143 MB/s | 0.84 |
JPEG 2000 is the slowest codec. rust-j2k is used because it agrees with OpenJPEG to within 1 grey level.
Reproduce
# micro-benchmarks (about 5 minutes; corpus files are used when present)OPENREADOUT_CORPUS_DIR=$PWD/corpus/files cargo bench -p openreadout-benchcargo nextest run -p openreadout-bench # memory ceiling and synthetic-input tests
# comparison with Bio-Formats and bioio (needs the oracle environment and oracle/bftools)cargo build --release -p openreadoutcargo xtask corpus fetch --tier smoke # plus the standard and full files named in bench/compare.pyuv run --project oracle python bench/compare.py run --reps 5 --max-load 3uv run --project oracle python bench/compare.py report
# Python arrays (build the Python extension first)oracle/.venv/bin/python bench/python_arrays.py run --reps 3oracle/.venv/bin/python bench/python_arrays.py report
# codecs and ND2 plane readscargo run --release -p openreadout-codecs --example codecbench -- <codec> <dir>cargo run --release -p openreadout-nd2 --example nd2_throughput -- FILE.nd2
# profiling build: release code with symbols, for samply, cargo-flamegraph or macOS `sample`cargo build --profile profiling -p openreadoutKnown limits
- All numbers come from one machine under shared load. A run on a quiet machine and on Linux x86-64 would make them firmer.
infoon mzML reads every spectrum header, so its time grows with the number of spectra. JCAMP-DX files are parsed whole.