Inside the files instruments write
A slide-scanner file is less like one long picture and more like a warehouse with a catalogue, and each instrument maker designed its own. This page looks inside a few of them and shows why OpenReadout can answer questions about a four-gigabyte file by reading a few megabytes of it.
A file is a warehouse with a catalogue
The Zeiss file on the home page is 3.7 gigabytes. It holds shelves of small tiles of pixels, a description of how everything was recorded, and signposts that say where everything is: a header at the very start and a catalogue of every tile at the very end. The rest is leftover space.
Here is that layout, drawn to scale from the file’s own records. Each of the 3.7 billion bytes has a place on the bar. The two strips below it magnify the first 800 kilobytes and the last 12.5 megabytes, where the signposts live.
Reading only what a question needs
The catalogue records where each tile sits in the picture and in the file, so a program that reads it doesn't have to load the whole file. OpenReadout reads the header, the description and the catalogue first, 1.4 megabytes in all. Then it fetches only the tiles that cover the part of the picture it was asked about, at the level of detail it was asked for. Pick a view and watch which bytes it needs.
Whichever view you pick, OpenReadout reads between 2.2 and 3.1 megabytes, less than a thousandth of the file. That's why info can describe a
multi-gigabyte slide in a fraction of a second, and why an agent can look closely at part of a file without loading all of it. It also helps with files
on a network share or in object storage, since OpenReadout fetches only the blocks it needs.
How an agent looks at a picture
Agents that browse the web look at screenshots. preview does the same for instrument data: it renders any image as a small picture with
rulers in full-resolution pixels and a scale bar, so the agent can read coordinates off the rulers and ask for a closer look.
Here are the three pictures OpenReadout returns on the way to a practical question: is this scan sharp enough to count cells, or does the slide need to be scanned again? The notes between them are the kind of reasoning an agent does with each one.
Every instrument keeps its own kind of catalogue
Zeiss’s warehouse is only one design. Thermo Fisher stores a mass spectrometer run as one long stream of scans with an index at the end. Flow cytometers write three blocks, the middle one a list of key–value pairs. Gatan’s electron-microscope files are a single tree of named tags that contains everything, the picture included. Each maker designed its layout on its own, often decades ago, and many have changed it between software versions.
Reading without the maker’s code
We worked out each layout from public sample files, using hex dumps and comparing our results with what independent open-source readers report. Our clean-room policy rules out vendor SDKs, headers and DLLs, and GPL reader source. Each step goes in the format’s provenance log. We then test each reader pixel for pixel against independent readers on public files, and each result carries an assurance verdict that says whether files like it were part of those tests.
A whole lab at once
Real questions are rarely about one file. A lab’s shared drive holds years of runs from every instrument it has owned, under names like
run_007_final2.nd2. Before anyone can analyse them, someone has to find the right ones and make sure they are intact.
So OpenReadout also catalogues whole folders. We pointed it at the collection of public files this project is tested on: . It read every header, checked every structure and wrote an index in . After that, questions that would have meant opening a thousand files in a dozen programs come back at once. Try one:
The same index produces a health report for the whole drive. In this collection it found
Once the right files are found, one command measures all of them and returns a single table, joined to the lab’s own sample sheet if there is one. Here, spike analysis of every sweep in recordings took a twenty-fifth of a second. An excerpt, from two very different neurons:
The same works for chromatogram peaks, plate-reader curves, qPCR runs, NMR spectra, cytometry gates and image statistics, across as many files as the lab has. Index and search a lab share shows how.
About the files. The figures use public files: Zenodo records, OME sample data and the test data of open-source projects such as pyABF and FlowKit, used under their licences. The file list and licences are in corpus/manifest.toml. The anatomy is described only as far as the images show it.