Sciex .wiff + .wiff.scan
Sciex Analyst writes an acquisition as a .wiff file (method, samples, scan index) and a .wiff.scan file (the scan data). OpenReadout returns MRM data and TOF data-dependent acquisitions as spectra, plus TIC, BPC and LC device traces. Q1/Q3 scans, enhanced product-ion, MRM³, SWATH and precursor/neutral-loss scans are not decoded (refused). Confidence: low.
Derived from public MetaboLights files (QTRAP 6500 and 6500+ MRM, TripleTOF 6600 DDA, an older ten-sample MRM file) by hex dump and comparison with the depositors’ own conversions and with ProteoWizard conversions made with the vendor library. The container is an OLE2 compound file read with openreadout-core::cfb, written from Microsoft’s public [MS-CFB] specification. No Sciex library, document or converter source was used. Details: docs/provenance/sciex-wiff.md. Format id: sciex-wiff; crate openreadout-sciex.
.wiff2 is not read. SCIEX OS .wiff2 files are encrypted; the reader recognises the extension (detection definite, with a note) and open returns exit 6 with a hint to export mzML from the vendor software. No attempt is made to read their content (docs/legal/clean-room-policy.md, rule 4).
A legacy acquisition is two files with one stem: <name>.wiff (method, sample information, scan index) and <name>.wiff.scan (the scan data). The reader accepts either path (wiff_path); each file finds the other by name, case-insensitively. Without the .wiff.scan, info, info --view structure and vendor work (metadata and the index), spectra do not — unless the file uses the single-file layout of older Analyst QS files (QSTAR): each sample’s scans are then in a SampleSubtree/SampleN/Scan stream inside the .wiff, a 32-byte stream preamble followed by the same scan data as a .wiff.scan holds after its 44-byte header; index offsets count from the end of that preamble (STREAM_PREAMBLE). Such a file needs no companion (note single-file layout; assurance layout feature single file (Scan stream)). Detection: extension .wiff plus the compound-file signature (definite), .wiff alone (extension only), .wiff.scan (likely, opens the .wiff beside it), .wiff2 (definite, refused).
All integers and floats little-endian. Every stream below starts with a 32-byte preamble (STREAM_PREAMBLE: u32 0, u32 version 4 or 5, 24 zero bytes).
Streams of the .wiff
| stream | layout | our names | status |
|---|---|---|---|
SampleSubtree/SampleN/Idx |
after the preamble, 54-byte records (INDEX_RECORD), one per scan: u32 offset and u32 byte length of the scan’s data in .wiff.scan, counted from byte 0x2C of that file (SCAN_FILE_HEADER); f64 time in ms; u16 (6); f64 total ion current; f64 base-peak intensity; f64 (0); f64 base-peak position (a TDC bin in TOF data); u32 (0) |
IndexRecord, offset, byte_len, time_ms, tic, base_peak_intensity, base_peak_x, parse_index, index_trailing |
validated: times and TIC equal the exports’ |
MethodSubtree/Method1/DeviceMethod0/Period0/ExperimentK/ExperimentHeader |
u16 at 0x56 = polarity (1 negative, 0 positive); u16 at 0x7A = scan type (4 MRM, 8 TOF MS, 9 TOF product ion); u32 at 0xB0 = number of mass ranges | ExperimentHeader, scan_type, polarity, range_count, parse_experiment_header, SCAN_TYPE_MRM, SCAN_TYPE_TOF_MS, SCAN_TYPE_TOF_PRODUCT |
inferred (polarity by comparing a negative- and a positive-mode file of one study) |
…/ExperimentK/MassRangeEx/MassRangeEx (older files: …/MassRange/MassRange, u32 1 instead of 2 and the dwell time in ms in place of the expected retention time, kept as parameter dwell_ms) |
u32 2, u32 parameters per range, then per transition / mass range: f32 first mass (MRM Q1), u32, f32 second mass (MRM Q3), f32 expected retention time in minutes (scheduled MRM), u32, a u16 byte length + UTF-16 compound name, then per parameter a u16 byte length + UTF-16 key (DP, EP, CE, CXP, CES, IRD, …), f32 value, f32 value, u32 |
MassRange, first_mz, second_mz, expected_rt_min, name, parameters, parameter, parse_mass_ranges |
Q1/Q3: validated (the exports’ targets); the rest inferred |
SampleSubtree/SampleN/SampleDAM/sMRMPro_adw1/sMRMPro_adw_Times |
(u32 start ms, u32 end ms) per transition: the scheduled window | parse_windows, windows |
validated (see below) |
SampleSubtree/SampleN/TDCInfo |
f64 at 0x20: the TDC bin width in ns (0.025) | tdc_width_ns |
validated via the calibration |
SampleSubtree/SampleN/TOFCalibrationData |
default f64 a, f64 t₀ at 0x20; u32 0; u32 number of index records; from 0x38 (TOF_CAL_FIRST) one 20-byte record (TOF_CAL_RECORD) per index record: f64 a, f64 t₀, u32 |
tof_calibration, tof_default_calibration, tof_mz |
validated: base-peak m/z equal the exports’ to 1e-12 |
SampleSubtree/SampleN/DDERealTimeData |
one 32-byte slot (PRECURSOR_SLOT) per cycle and product-ion experiment, from the preamble on: f64 precursor m/z, f64, f64 (an integer; kept as precursor_slot_value), f64 |
precursor_slot |
validated: every precursor equals the export’s |
SampleSubtree/SampleN/Log |
UTF-16 text; key: value fields separated by commas: Component ID (the model), Serial Number, Manufacturer, Model, Firmware Version, … |
log, log_fields, utf16_runs |
read |
SampleSubtree/SampleN/SampleDABE/DATA |
after the preamble (its last u32 is 4 or 6 here) a u32 100, then u16-byte-length UTF-16 strings: sample name, sample id, comment, data file, acquisition method, … | strings, sample_strings |
read |
SampleSubtree/SampleTable |
u32 Unix time at 0x3E: the acquisition start (sample 1) | started_at |
validated against the export’s startTimeStamp |
SampleSubtree/SampleN/Devices/Device_K/Channel, DevData |
LC devices recorded with the sample: Channel after the preamble u32 1, u16 channel count, then per channel u32 1, u16-length UTF-16 name, u16-length UTF-16 unit ((psi)), f64 samples per second, u16 channel index, f32, f32, i32 (DeviceChannel, rate_hz, unit, parse_device_channels); DevData after the preamble one f64 per channel per sample (parse_device_data) |
validated: every value of six channels equals the export’s | |
FileRec_Str |
UTF-16 text naming the software (Analyst 1.7.2, Analyst TF 1.8.1) |
read | |
MethodSubtree/Method1/AcqMethodFileInfoStm |
UTF-16 path of the acquisition method (.dam) |
read | |
| everything else | listed by info --view structure with its size; vendor lists every stream |
– |
.wiff.scan
A 0x2C-byte header, then each scan’s data where the index points.
- MRM (scan type 4;
expand_zero_runs): f32 values in which a negative value −(n + 0.01) stands for n zeros. A cycle expands to 2 × (number of transitions) values whose second half holds the transitions in method order (QTRAP 6500/6500+; the first half is zero in every cycle of the corpus files and is not interpreted), or (older Analyst files) to exactly one value per transition. Any other width is an error. - TOF (scan types 8 and 9;
decode_tdc): a byte stream of the time-to-digital histogram in steps oftdc_stepTDC bins (u32 at 0x38 of the experiment’sExperimentHeaderEx: 4 in the TOF experiments of the corpus).FF FF FF FF+ u32 sets the TDC bin; a byte 0x00–0x7B is a count (then the next step); 0x7C, 0x7D, 0x7E are followed by a 1-, 2- or 4-byte count; 0x80–0xFB skip (byte − 0x80) empty steps; 0xFC, 0xFD, 0xFE are followed by a 1-, 2- or 4-byte number of steps to skip; 0xFF bytes running to the end of a scan end it; 0x7F is not known and is an error. Validated on every scan of the corpus TOF files: the counts sum exactly to the index TIC, and the index’s base-peak position is the decoded profile’s first maximum. m/z of TDC bin b: (a·(width·b − t₀))² with the scan’s (a, t₀) and the TDC width.
Mapping
- One run per sample (
Sample,number;spectra[k].index= sample − 1). Index record i belongs to cycle ⌊i / E⌋ + 1 and experiment (i mod E) + 1 of the E experiments of period 0 (Experiment,number,header,ranges); records of length 0 (an experiment that recorded nothing in that cycle) give no spectrum. - MRM: one spectrum per cycle and precursor (Q1): the transitions scheduled at the cycle’s time (window start ≤ t ≤ end; without windows, all), m/z = Q3, sorted;
precursor_mz= Q1;collision_energy= the transitions’CEwhen they share one (else per transition inextra.collision_energies);activationCID;ms_level2;centroided;native_idsample=S period=1 cycle=C experiment=E transition=K(K the group’s first transition);extra:cycle,experiment,scan_type,index_record,transitions,compounds. Scheduling exception (inferred from one file, 17 of 165 windows): a scan at a window’s start or end time whose whole cycle is zero (TIC 0) — written while the instrument switches transitions — is left out of the transitions starting or ending there (not the first scan of the run). - TOF: one spectrum per non-empty record: the histogram as a profile (non-empty steps plus the empty step on each side of a run),
centroidedfalse,ms_level1 (TOF MS) or 2 (product ion);total_ion_current= index TIC;base_peak_mz/base_peak_intensityfrom the index (for a scan that stores no position: the first maximum of the profile); product ions:precursor_mzfromDDERealTimeData,collision_energyfrom the experiment’sCEparameter,activationHCD(beam-type CID) —CIDwhen the writer is Analyst QS (the QSTAR exports’ label; no field states it, inferred from the exports);native_idsample=S period=1 cycle=C experiment=Eas in the exports;extra.calibration(a, t₀, bin width). scan_number= position + 1 (MRM groups share an index record).traces[]: per sampleTIC(index TIC) andBPC; after them one trace per LC device of each sample (LC device K: Column Pressure, Pump A Flowrate, …: one channel per device channel with its unit,sample_rate_hzfrom the channel records, time from 0;extra.sample,extra.device) (index base-peak intensity; for MRM, which stores none, the largest transition value of the cycle), one point per non-empty record, irregular sampling.- Run
extra:sample_name,sample_id,acquired_at(SampleTable+ 0x3E: the acquisition computer’s local clock as seconds since 1970; written with the UTC offset it has from the compound file’s earliest storage creation time, a UTC FILETIME, rounded to a quarter hour, else with no zone),instrument_serial,method,polarities,experiments(scan type, polarity, mass-range count per experiment),cycles,stored_spectra(SRMorprofile),method_summary.instrument: manufacturerSCIEX, model =Component ID, software and version fromFileRec_Str.
check finding codes
truncated (index ends inside a record; a scan past the end of .wiff.scan), container (compound-file problems), undecodable (errors); missing_file (no .wiff.scan), missing_method, time_not_monotonic, window_count_mismatch, tic_mismatch (warnings).
Vocabulary (every public identifier in crates/openreadout-sciex must appear here)
| identifier | meaning |
|---|---|
SciexWiffReader, SciexDataset, FORMAT_ID, open, open_input, wiff_path |
reader, opened file, the id sciex-wiff, open (a local path, or an input whose namespace holds the .wiff and .wiff.scan), the .wiff a path names |
STREAM_PREAMBLE, INDEX_RECORD, SCAN_FILE_HEADER, TOF_CAL_RECORD, TOF_CAL_FIRST, PRECURSOR_SLOT |
sizes and offsets above |
SCAN_TYPE_MRM, SCAN_TYPE_TOF_MS, SCAN_TYPE_TOF_PRODUCT |
scan type codes 4, 8, 9 |
le_u16, le_u32, le_f32, le_f64 |
bounds-checked little-endian readers |
IndexRecord, offset, byte_len, time_ms, tic, base_peak_intensity, base_peak_x, parse_index, index_trailing |
the scan index |
expand_zero_runs, decode_tdc |
MRM and TOF scan decoders |
MassRange, first_mz, second_mz, expected_rt_min, name, parameters, parameter, parse_mass_ranges |
transitions / mass ranges |
ExperimentHeader, scan_type, polarity, range_count, parse_experiment_header |
experiment header fields |
Experiment, number, header, ranges, tdc_step |
one experiment of the method; TDC bins per stored TOF step |
Sample, index, windows, log, strings, started_at |
one sample (run): scan index, scheduled windows, log text, sample strings, acquisition start |
parse_windows, tof_calibration, tof_default_calibration, tof_mz, precursor_slot, tdc_step |
scheduling windows, TOF calibration, precursor slots |
utf16_runs, sample_strings, log_fields |
text helpers |
DeviceChannel, name, unit, rate_hz, parse_device_channels, parse_device_data, devices |
LC device channels of a sample and their values |
How this reader was derived, file by file: provenance log.