Many files: batch tables and sample sheets
Most questions span a folder: mean intensity per condition, median CD4 per donor, which wells failed. OpenReadout answers them in one command. It measures every file, joins the result to your sample sheet or plate layout, and summarizes it by group.
Here are four images, two per condition, and a sample sheet that names them:
file,conditionwell_A1.nd2,DMSOwell_A2.nd2,DMSOwell_B1.czi,drugwell_B2.czi,drugopenreadout stats plate/ --sample-sheet samples.csv --by condition4 data sets (4 ok, 0 failed)summary by condition, channel_name of mean, median (4 rows used, 0 left out)condition channel_name value n mean sd sem median min max cv_percent───────── ──────────── ────── ─ ───── ── ─── ────── ───── ───── ──────────DMSO DAPI mean 2 12.75 0 0 12.75 12.75 12.75 0DMSO DAPI median 2 8.5 0 0 8.5 8.5 8.5 0drug c0 mean 2 35 0 0 35 35 35 0drug c0 median 2 35 0 0 35 35 35 0samples.csv joined on `file` = path: 4/4 rows, 4/4 data sets annotated (added: condition)OpenReadout worked out the join key and reported it on the last line: the sheet’s file column matches the file names. It also added the channel name as a group column, because the two conditions were imaged with differently named channels and it won’t average across channels.
Several files at once
info, check, stats and export accept several files, directories and glob patterns:
openreadout check -r /data/2026-09-21 # one report per file, then a summaryopenreadout info -r --jsonl --skip-unknown /data > info.jsonl # one JSON object per lineopenreadout info --json "/data/**/*.nd2" # quoted: openreadout expands the globopenreadout export -r --skip-unknown raw/ -o ome/ # mirrors raw/ under ome/- Directories are walked in name order.
-ralso walks sub-directories. Hidden entries are skipped. - A directory that is itself a data set, such as a ChemStation
.D, a Waters.rawor a Bruker.d, is one input. - A failed file does not stop the run.
--fail-faststops at the first failure. - The exit code is the worst code of any input, so
check -r DIRexits 4 when any file is corrupt. See Exit codes. - Batch tables are the exception:
batch, and any command that builds a table (--tidy,--sample-sheet,--by,--csv, or several inputs totrace,tableandgate), exit 0 once the table is built. A data set that fails is a row with anerrorcolumn and is counted in theN data sets (N ok, N failed)line. Pass--fail-fastto stop at the first failure and exit with its code instead. --skip-unknownleaves out files that are not instrument data, such as READMEs and spreadsheets.
path format contents status─────────────────────────── ─────────── ─────────── ───────────/data/2026-09-21/a.nd2 nd2 intact ok/data/2026-09-21/cut.czi czi 4 findings corrupt (4)/data/2026-09-21/run.D chemstation intact ok3 inputs: 2 ok, 1 failedEach line of the --jsonl output is one JSON wrapper with a path field. It loads into pandas directly:
import json, pandas as pdrows = [json.loads(line) for line in open("info.jsonl")]df = pd.json_normalize([ {"path": r["path"], **img} for r in rows if r["ok"] for img in r["data"]["images"]])One table from many files
info, stats, trace, table and analyze gate can return one tidy table for all inputs instead of one report per file. Ask for it with --tidy. Any of --sample-sheet, --by, --where, --fields, -o, --csv or --from-index turns it on as well. For trace, table and analyze gate, several inputs, a directory or a glob is enough.
# median FITC-A per conditionopenreadout table runs/ --sample-sheet samples.csv --where parameter=FITC-A --by condition --value median
# mean intensity per condition and channel; wells averaged first; Welch's t-test against DMSOopenreadout stats plate/ -r --sample-sheet layout.xlsx --by condition --replicate well \ --test welch --control DMSO -o per_well.parquet
# population counts and median fluorescence per FCS fileopenreadout analyze gate fcs/*.fcs --workspace analysis.wsp --median Comp-FITC-A -o populations.csv
# header metadata of a whole share, one row per data setopenreadout info -r share/ --tidy --fields sample_id,instrument_model,started_at --csvRows and columns
Each row is one data set times the measure’s grain:
| command | one row per | values |
|---|---|---|
stats |
image × channel; --per image, plane, well or field changes it |
mean, median, min, max, std, percentiles, saturated fraction |
trace |
trace × sweep × channel | mean, max, min, std, duration_s |
table |
FCS parameter, or plate well × read | FCS: median, mean, sd; plates: value |
analyze gate |
population | count, percent_of_parent, median:<parameter> |
info |
data set | the columns of the index |
Columns come in a fixed order. path, format and files are first. Then come the keys (image, channel, parameter, well), what the file records about the sample, your sample sheet’s columns, the values, and finally error and error_code. Names are lower case and stable. A unit that never changes is part of the name (duration_s, wavelength_nm). The JSON output lists every column with its type, unit and description.
A file that cannot be read becomes a row with error and error_code set, and the run carries on. Files that belong together, such as a multi-file OME-TIFF or a SpikeGLX .meta and .bin, are measured once, and the files column counts them. Your sample sheet and the output file are not treated as inputs.
Narrow the table with these flags:
--where COLUMN=VALUE(or!=, repeatable) keeps matching rows.--fieldskeeps only some columns.--formats fcs,czikeeps only those formats.--from-index INDEX --query "..."takes the inputs from a lab-share index.
Output
- No flag: a text table of the first 40 rows, the join report and notes.
--csv: CSV on standard output.--jsonl: one object per row and line.--json: the JSON wrapper withcolumns,rows,summaryandjoins.--limitand--offsetpage the rows.-o FILE:.csv,.tsv,.jsonl,.jsonor.parquet. With--by, the file holds the summary. Parquet keeps each column’s unit and description.
-o writes to a temporary name, reads the file back, then renames it into place. While it runs it keeps a journal next to the output. If a run is interrupted, run the same command again: files that have not changed are not measured a second time.
Analyses over a folder
openreadout batch MEASURE INPUTS… runs any measure over many files with the same table flags. The measures are stats, trace, table, info, spectra and the analyses peaks, chromatogram, assay, nmr-peaks, ephys-features, spikes, qpcr and gate.
Pass options by name with --set KEY=VALUE, or all at once with --options '{...}'. A value is read as JSON when it parses. The option names are those of the matching MCP tool. The rows option picks which records become rows: for example peak, compound or chromatogram for peaks, and sweep, cell or spike for ephys-features. Each row holds the same numbers as the single-file analyze output.
openreadout batch peaks runs/ -r --set rows=chromatogram --sample-sheet samples.csv --by conditionopenreadout batch ephys-features cells/ -r --set rows=cell --sample-sheet genotypes.csv --by genotypeopenreadout batch qpcr plates/ --set ddcq=true --set 'reference_targets=["18s"]' --by treatmentopenreadout batch stats plate/ --set per=well --sample-sheet layout.csvThe analyses themselves are described in Recipes.
Sample sheets and plate layouts
--sample-sheet FILE (also --samples or --layout, repeatable) reads two kinds of file.
A sample sheet is a table with a header row. It can be CSV, TSV or another delimited text file, or a worksheet of an XLSX, XLS, XLSB or ODS workbook. --worksheet NAME picks the worksheet. A sheet with Row and Column number columns gets a well column.
A plate layout (plate map) holds one grid per annotation. The header row numbers the plate columns from 1, and each following row starts with a plate-row letter. The grid’s top-left cell names the annotation; if it is empty, the line above the grid or the worksheet name does. Several grids can sit one under another or on separate worksheets. 96-, 384- and 1536-well plates are read.
condition,1,2,3A,DMSO,DMSO,drugB,DMSO,drug,drugHow rows are matched
OpenReadout tries each sheet column against several keys of every data set:
- the path, file name or file stem
- the sample id, name, barcode or vial the file records (see the experiment block)
- the plate well (
A1,A01andR1C1are the same well) - the run order, from the acquisition start time
The pair that matches the most rows wins. When rows still match several sheet rows, a second column joins the key; plate barcode plus well is the usual case. Sheet values with * or ? are patterns (ctrl_*.fcs). To set the key yourself, use --key SHEET_COLUMN=FIELD, and repeat it for a composite key.
Unmatched rows are reported, not dropped. The join report lists:
- the key chosen and the runners-up
- the columns added
- data sets that no sheet row matched
- sheet rows that matched no data set, which is often a misnamed file
- rows that matched conflicting sheet rows; these are left without annotations
A sheet that matches nothing is a usage error (exit 2). Its message shows the sheet’s columns and the first file names.
Sheet columns whose values are all numbers are typed as numbers, so dose sorts numerically.
Wells of high-content screening images need no extra work. stats rows carry the well when the reader records one, and a plate layout joins onto it.
Group summaries
--by condition[,dose] adds one summary row per group and value column. Each row has n, mean, sd, sem, median, min, max and cv_percent. --value picks the value columns; the default is the measure’s main values.
Different measurements are not pooled. Columns that say what was measured, such as channel, FCS parameter, population, trace or wavelength, are added to the groups whenever they vary. --exact-by turns this off. Images, wells, sweeps and files are pooled.
--replicate COLUMNaverages each replicate (a well, a biological replicate) first, soncounts replicates.--test welchor--test mann-whitney, with--control VALUE, compares every group with the control group. It addsdiff,ratio,statistic,dfand a two-sidedp_value. The tests match SciPy’sttest_ind(equal_var=False)andmannwhitneyu. P-values are not corrected for multiple comparisons.
Rows with an error, rows filtered out and rows without a value are left out and counted.
To regroup a table you saved earlier (CSV, TSV, JSON Lines, JSON or Parquet), use batch summarize:
openreadout batch summarize per_well.parquet --by condition,dose --value meanFinding the same sample across instruments
openreadout link DIR… reads the headers of every data set and groups the files that measured the same sample. Typical pairs are a Thermo .raw file and its mzML conversion, or two instruments’ files with the same plate barcode.
openreadout link share/ --jsonEach link lists its evidence and a confidence:
- High: the same barcode; the same plate barcode and well; the same recorded sample id; one file names the other as its source; the same acquisition start, run length and instrument; the same sequence of mass spectra.
- Medium: the same instrument serial, scan count and run length; the same sample name where no id is recorded; the same file stem in two formats.
- Medium or low: one file’s sample id appears in the other’s file name.
A matching well on its own doesn’t link two files, because every plate has the same wells. Identifiers that name controls (blank, QC, pool, standard) count as weak evidence. So do identifiers that more than --max-shared data sets share (default 12). Links at or above --min-confidence (default medium) form groups; weaker ones are listed separately. A group whose files record different sample ids is reported as a conflict.
Shell loops
Branch on the exit code, not on the text:
for f in /data/*; do openreadout check "$f" --json > "reports/$(basename "$f").json" case $? in 0) ;; # intact 3) rm "reports/$(basename "$f").json" ;; # not an instrument file 4) echo "CORRUPT $f" ;; 6) echo "cannot read yet: $f" ;; *) echo "failed: $f" ;; esacdoneThe repository has three ready-made scripts in examples/shell/. They need jq and use $OPENREADOUT as the binary when it is set.
batch_summarize.sh DIRprints one tab-separated line per file: format, image count, dimensions, pixel type, pixel size and acquisition time.triage_corruption.sh DIRprintsOK,CORRUPTorUNSUPPORTEDper file with the findings, and exits 1 when anything is damaged. Run it on a copy before deleting the original from the acquisition PC.export_planes.sh OUT FILES…exports one plane per image (channel 0, middle z, first time point) and checks that each export was verified.CHANNEL=1 TO=ome-zarrchanges the channel and the format.
From Python, openreadout.batch() (with "summarize" for batch summarize) and openreadout.link() take the same options as the commands. See the Python guide.
Validation
The table measures, sample-sheet reading, group statistics and link are checked against independent readers and SciPy on the test corpus. The method and results are on the validation page.