Agilent OpenLab CDS injections (.dx, .rx)
Agilent OpenLab CDS 2.x writes one .dx container per injection, with the data-analysis results in a .rx file of the same name. OpenReadout returns the detector signals of an injection and the integrated peaks of its results. Spectra directories, MS data, calibration and compound amounts are not read.
Derived from public files (Zenodo record 14316687, CC-BY-4.0, and allotropy’s MIT-licensed test data) and allotropy’s decoder (MIT, read as documentation). rainbow-api (LGPL), chromConverter (GPL) and allotropy are run as reference readers only. See docs/provenance/openlab-cds.md. Format id openlab-cds; crate openreadout-chrom (openlab_*.rs; zip containers through openreadout_core::zip). Confidence: medium.
OpenLab CDS 2.x keeps a sequence run as a result set: a folder <name>.rslt holding one injection container <injection>.dx per injection, the data-analysis results of each injection in a .rx package with the same stem, the sequence file <name>.acaml, and method files (.amx acquisition, .pmx processing, .sqx sequence template, .scml, .mfx). OpenReadout opens an injection: the .dx (or its .rx, which opens the .dx of the same name when present). A .rslt folder is refused with a hint listing its .dx files (openreadout batch 'set.rslt/*.dx' reads them all).
Containers
.dx and .rx are zip archives (PKWARE APPNOTE; members stored or deflated) laid out as Open Packaging Conventions packages: [Content_Types].xml maps extensions to content types and _rels/.rels points at the manifest. looks_like_openlab recognises the first member (injection.acmd, Base/InjectionACAML, or a GUID-named part such as b11988f5-7a62-4da7-903a-17df3a95267d.CH): a definite detection with a .dx/.rx extension, a likely one without. A first [Content_Types].xml or _rels/.rels (any Open Packaging Conventions file, e.g. .xlsx) counts only with a .dx/.rx extension; any other zip with our extension is extension-only. JCAMP-DX files share the .dx extension but are text; they never start with PK\x03\x04.
| member | content type (from [Content_Types].xml) |
what it holds |
|---|---|---|
injection.acmd |
Agilent.OpenLab.Rawdata/ACMD |
the injection manifest (XML) |
<trace id>.CH |
Agilent.OpenLab.Rawdata/Signal179 |
one detector signal |
<trace id>.IT |
Agilent.OpenLab.Rawdata/InstrumentTrace179 |
one instrument curve (pressure, flow, solvent ratio, temperature, lamp voltage, …) |
<trace id>.UV |
Agilent.OpenLab.Rawdata/Spectra131 |
DAD spectra: a ChemStation-style version-131 .uv file (read with the ChemStation reader) |
<trace id>.UVD, _rels/<trace id>.UV.rels |
spectra directory and its relationship | listed, not decoded (an index of the .UV part) |
| other parts | other types | listed, not decoded |
Base/InjectionACAML (in .rx) |
text/xml |
the injection’s integration results |
Base/AuditTrail (in .rx) |
application/octet-stream |
binary audit trail with readable user/computer/message strings; listed only |
Members are read whole into memory (limit 512 MiB each), checked against their CRC-32; encrypted members are refused (exit 6).
Manifest (injection.acmd)
UTF-8 XML (with BOM), namespace urn:schemas-agilent-com:acmd20, one InjectionInfo. Elements are matched by local name.
| element | our name (InjectionManifest) |
example |
|---|---|---|
Version |
version |
2 |
Location |
location → extra.vial |
71, 1, D1F-A8 (a string: plate positions occur) |
InjectionSource |
injection_source |
GC Injector, HipAls |
InjectionVolume, InjectionVolumeUnits |
injection_volume, injection_volume_unit → extra.injection_volume (µL; nL and mL converted) |
1 μL (U+03BC), 5 µL (U+00B5) |
SequenceLine, Replicate |
sequence_line, replicate |
2, 1 |
SampleName, RunOperator, Barcode |
sample_name, operator, barcode |
FKB-FA-035-A1-F, Admin |
RunDateTime |
run_started → extra.acquired_at |
2023-02-23T15:05:20.4882635+01:00 (start of the run; the vendor CSV’s “Injection date” is the end) |
AcquisitionMethod |
acquisition_method → extra.method_path, extra.method (file name) |
D:\Data\GC-FID\Results\FKB\FKB-FA-035-08.rslt\Polyarc_90_2_30_280_3.amx |
Signals/Signal |
signals[] (ManifestSignal) |
one per part |
ManifestSignal: Encoding → encoding (content type; kind() is its last component, Signal179), TraceId → trace_id (names the part), DeviceName → device (FID, DAD, PMP, THM, WPS, FLD), DeviceNumber → device_number, ChannelName → channel (FID2B, DAD1D, PMP1A), Description → description (DAD1D,Sig=228,4 Ref=360,100, PMP1A,Pressure), TimeStart/TimeEnd → time_start/time_end (ms for detector signals = the part header’s first/last time; for instrument curves 0 and the value count), Minimum/Maximum → minimum/maximum, Slope → slope (= the part’s scale factor), NumberOfValues → declared_values, Units → units, IsIntegrable → integrable (true for detector signals). A Signal without TraceId is skipped. Empty GenericResults… elements are ignored.
Parts
Every decoded part is a ChemStation version-179 file (docs/formats/chemstation.md): 03 31 37 39 at byte 0, file type OL DATA FILE, the 0x1800 UTF-16 header (sample 0x35A, operator 0x758, date 0x957, method 0xA0E, units 0x104C, signal description 0x1075, scale f64 at 0x127C). The ChemStation header parser and body decoders are reused unchanged.
- Detector signal (
Signal179) — body: little-endian f64 values (Float64;SecondDifferenceandDeltaRecordsbodies are decoded the same way if a part carries another version).value = raw × scale. Time axis: the header’s first and last time (f32 ms at 0x11A/0x11E) overn − 1intervals, as for ChemStation version 179. In the GC corpus the first time is 19.812 ms and the samples are 20 ms apart; the header’s last time is 680,000 ms = 34,000 × 20 ms, so this rule gives 20.0000055 ms (0.19 ms, 3·10⁻⁶ min, of drift at the end of an 11 min run); both oracles use the same grid, and the vendor’s peak limits fall on it to 10⁻¹² min. - DAD spectra (
Spectra131) — the part is a ChemStation version-131.uvfile (docs/formats/chemstation.md), opened in memory with the ChemStation reader: one trace, one channel per wavelength, one sample per spectrum. Its records carry tag 70: after the 22-byte record header (tag, length, time in ms, first/last/step wavelength × 20) little-endian f64 values, one per wavelength, instead of tag 67’s 16-bit differences.value = stored × scale × scale(mAU): the header’s scale factor (0.000476837158203125 = 1/2097.152, equal to the manifest’sScaleFactor) applies twice. rainbow-api’sparse_uvreturns stored × scale; the second factor is established against the DAD’s own channels of the same injections (below). The manifest’sNumberOfRecordsis the spectrum count (declared_records). - Instrument curve (
InstrumentTrace179) — body: 16-byte samples (PAIR_BYTES), each a little-endian f64 time in ms and a little-endian f64 raw value;value = raw × scale(1000 × 0.001= 1.0 mL/min flow;380 × 0.1= 38 % solvent B). The header’s time fields are not a time range for these parts. When the times are evenly spaced (all corpus curves: 60, 97.875, 100 and 1,000 ms apart) the trace is regular (extra.axiswithquantitytime); otherwise it gets a second channeltime(minutes) andsample_rate_hz0 (extra.irregular_times). A body cut inside a sample ends withBodyEnd::Truncated.
Results (.rx → Base/InjectionACAML)
XML, namespace urn:schemas-agilent-com:acaml21 (schema version 2.1.30.999). Read into InjectionResults: Doc/DocInfo/CreationDate → processed_at; Content/Method/Name, OriginalVersion, QuantitationMethod → processing_method, processing_method_version, quantitation (ESTD); per Injections/Result: InjectionMeasData_ID/@id → measurement_id, DataAnalysisSoftware (else DataAnalysisApplication/AgilentApp name and version) → software, Integrator → integrator (TwelveTone), ProcessingStatus/TransformationChainState → processing_state (Passed), and one SignalResults per SignalResult (Signal_ID/@id → signal_id, Peaks → peaks).
Peak element |
VendorPeak field |
unit |
|---|---|---|
@id |
peak_id |
|
RetentionTime |
rt_min (a peak without a finite one is skipped) |
min (s, ms, h converted) |
Type |
peak_type |
NormalPeak |
Area |
area, area_unit |
as written: pA·s, mAU·s |
AreaPercent, HeightPercent |
area_percent, height_percent |
% |
Height |
height, height_unit |
pA, mAU |
Symmetry |
symmetry |
|
BeginTime, EndTime |
start_min, end_min |
min |
BaselineCode |
baseline_code (trailing blanks removed) |
BB |
BaselineStart, BaselineEnd |
baseline_start, baseline_end |
signal units |
WidthBase |
width_base_min |
min (the vendor CSV’s Width [min]) |
BaselineModel |
baseline_model |
Linear |
BaselineRetentionHeight |
baseline_at_apex |
signal units |
LevelStart/LevelEnd, PurityPassed, MSPurityPassed, BaselineParameters and custom fields are kept in info --view full (vendor.results), not in the table. Each peak’s InjectionCompound (linked by Identification/Qualified/Peaks/Peak_ID/@id) gives compound (CompoundName, empty names are none; every corpus file has empty names), compound_type (Type) and expected_rt_min (ExpectedRetTime, finite values only; the corpus writes -INF).
Sequence file (<result set>.acaml)
Same namespace. Read into SequenceFile: Resources/Instrument/Name → instrument_name (Luxo HPLC, a configured name), Technique → technique (LiquidChromatography), Modules → modules (InstrumentModule: @id → id, Name → name, Manufacturer → manufacturer, Type → kind (Pump, Sampler, ColumnCompartment, Detector), PartNo → part_number (G7117C), SerialNo → serial_number, FirmwareRevision → firmware); per Injections/MeasData (SequenceInjection): @id → measurement_id (= the .rx’s InjectionMeasData_ID), BinaryData/DataItem/Name → data_file (the .dx name), Signals → signals (SequenceSignal: @id → id (= the .rx’s Signal_ID), Name → name, Description → description, Type → kind, TraceID → trace_id (= the manifest’s TraceId), DetectorName → detector, InstrumentModule_ID/@id → module_id); the first injection’s AcquisitionApplication/AgilentApp (else AcquisitionSoftware) → acquisition_software. injection(dx_name, measurement_id) finds this injection by file name, else by measurement id. Every .acaml in the .dx’s folder is tried (up to 64 MiB each); the one listing the injection is used.
Which signal a vendor peak belongs to (SignalMapping)
The .rx names signals by an id of the sequence file, not by the manifest’s TraceId:
SequenceFile— the folder’s.acamllists the injection:Signal_ID→SequenceSignal.trace_id→ our signal (tables[0].extra.signal_mapping=sequence_file).BaselineMatch— no sequence file (the Polyarc data set): the vendor’sBaselineStart/BaselineEndof a peak are the raw signal at the samples nearestBeginTime/EndTimefor baseline-to-baseline ends. For each result signal with peaks, the detector signal whose samples equal the most of these values (within 10⁻⁶ relative) is chosen; it must be a strict winner and different result signals must land on different signals (baseline_match). On all 54 GC injections this picks FID2B, the signal the vendor’s CSV export names.SingleSignal— one detector signal, one result signal with peaks.Unmapped— otherwise; the table’ssignalcolumn then holds the result’s signal id andtraceis NaN;checkreportsresults_unmapped.
Mapping to the data model
traces[]: detector signals in manifest order, then instrument curves in manifest order, then spectra parts (their trace from the ChemStation reader under the manifest description,extra.spectratrue,extra.wavelength_range_nm,extra.record_encodingfloat64) (soanalyze peaks FILEintegrates the first detector signal by default). Name = the manifest description (else the channel), one channel named afterChannelName(FID2B) with the manifest’s unit (else the header’s),dtypefloat64,scalefrom the header.extra:file,part,trace_id,encoding,format_version,channel,signal,detector,device_number,integrable,declared_values,sample_name,operator,acquired_at,method,method_path,vial,injection_volume(µL),injection_source,sequence_line,replicate,barcode,separation(GC/LCfrom the sequence file’s technique, elseGCwhen the injection source starts withGC),software(OpenLab CDS),software_version(sequence file),instrument_name(sequence file),instrument(the recording module’s part number),instrument_serial,module,module_firmware(the module the sequence file links to the signal),wavelength_nm,bandwidth_nm,reference_wavelength_nm,reference_bandwidth_nm(fromSig=/Ref=),x_start_min,x_end_min,axis,irregular_times,vendor_peak_count. The experiment model reads sample, vial (sequence position), operator, start, method, injection volume, separation, instrument model and serial fromtraces[0].tables[0]vendor_peaks(when a.rxwas read), one row per vendor peak, grouped by trace, in retention-time order. These values are calculated by the vendor software, not by OpenReadout (extra.source). Columns:signal(code intoextra.categories),rt_min,start_min,end_min,area(the vendor’s unit, e.g.pA·s—openreadout analyze peaksreports signal × min unless--area-seconds),height,area_percent,height_percent,width_base_min,symmetry,baseline_start,baseline_end,baseline_at_rt,baseline_code,peak_type,compound(codes intoextra.categories),trace(our trace index, NaN when unmapped).extra:source,results_file,software,processing_method,processing_method_version,quantitation,integrator,processing_state,processed_at,signal_mapping,peaks_by_signal.info --view full→vendor:content_types,parts(name, size, compressed size, method),manifestandresults(the XML converted without renaming),part_headers(the ChemStation header of every decoded part),results_file,sequence_file(instrument, technique, modules, injection count).info --view structure: every member (signal,instrument-curve,manifest,part),missing-partentries for manifest signals without a part, the.rx(results) and the.acaml(sequence).
check finding codes
bad_member (a member fails to inflate or its CRC-32), truncated, bad_record, bad_results (errors); the ChemStation reader’s findings on each spectra part, prefixed spectra <channel>:; missing_part, value_count_mismatch (part values ≠ NumberOfValues), record_count_mismatch (spectra ≠ NumberOfRecords), scale_mismatch (header scale ≠ Slope), no_samples, time_not_increasing, peak_outside_signal (warnings); not_decoded, results_unmapped (info).
Validation
Automated in cargo test -p openreadout-corpus-tests --features corpus --test corpus (traces, tables) and --test openlab (vendor results). Measured 2026-09-24:
| comparison | files | result |
|---|---|---|
detector signals against rainbow-api 1.5.2 (.CH parts extracted) and chromConverter 0.9.0 (whole .dx) — the two oracles agree exactly |
6 GC + 3 LC injections, 18 signals | bit-identical (xxh3 of every value), sample rates, first/last retention times, sample name, operator, method path |
| instrument curves against chromConverter | 3 LC injections, 29 curves | bit-identical values, times, units, names |
DAD spectra parts against rainbow-api 1.5.2 parse_uv (× the scale factor once more) — 2026-09-26 |
2 LC injections (chromConverterExtraTests MeOH1.dx, openlab.sirslt), 4,050 + 1,013 spectra × 156 wavelengths |
bit-identical (xxh3 of every wavelength channel), times, wavelengths |
DAD spectra averaged over each DAD channel’s band (Sig=210,4: 208–212 nm) against that channel, recorded separately by the detector — 2026-09-26 |
2 LC injections, 8 channels (200–360 nm) | slope 0.993–1.009 (1.023 at 200 nm, the edge of the range), r ≥ 0.99999: the unit of the spectra (a scale factor of 2,097 away from rainbow’s reading) |
detector signals and instrument curves of 3 more injections (MeOH1, a YADG GC injection, the .sirslt result) against rainbow-api and chromConverter — 2026-09-26 |
3 injections, 10 signals, 42 curves | bit-identical |
our vendor_peaks table against allotropy 0.1.146’s reading of the .rx |
54 GC + 12 LC injections, 482 peaks | 4,338 values equal bit for bit; every peak on the signal allotropy (via the sequence file) or the vendor CSV names |
the table against the vendor’s CSV export (_1.csv) |
54 GC injections, 196 peaks | 784 values within the CSV’s printed precision (RT 3 decimals, others 2), baseline codes equal; sample name, operator, vial, method, injection volume and the integrated signal equal (324 fields) |
our decoded signal integrated between the vendor’s limits above the vendor’s baseline (integrate_range), against the vendor’s area |
196 peaks | median 0.0006 %, p95 0.010 %, max 0.25 % |
| height above the vendor’s baseline (largest sample) against the vendor’s height (an interpolated apex) | 196 peaks | median 0.42 %, max 1.4 % |
openreadout analyze peaks --baseline drop (the default until 2026-09-24) against the vendor |
196 peaks | 196/196 found; retention time median 0.0001 min, max 0.004 min; area median 0.6 %, but p95 611 %: peaks on the solvent tail get the tail with a drop baseline |
the same with --baseline valley |
196 peaks | area median 0.19 %, p95 3.8 %, max 12.7 %; area % among the vendor’s peaks median 0.34, max 0.88 points |
the same with the default --baseline auto |
196 peaks | 196/196 found; area median 0.49 %, p95 4.0 %, max 35 %; area % among the vendor’s peaks median 0.09, p95 0.38, max 0.77 points |
OpenLab’s integrator ends these GC peaks where they meet the solvent tail (baseline-to-baseline on the tail): the default auto baseline does the same (baseline_reason: sloped_background; book/src/guides/quantitation.md → The auto baseline), and valley also matches here; or read the vendor’s own results from tables[0]. The LC result sets cannot validate integration: their publisher cut the detector signals to 4 values (check reports value_count_mismatch).
Performance (release build, Apple M-series, warm cache, measured 2026-09-24 with /usr/bin/time -l): info or check on a 560 KB GC .dx with its .rx under 5 ms and 13 MB peak memory; on an LC .dx with 16 parts, its .rx and a 391 KB sequence file 10 ms and 17 MB; analyze peaks --baseline valley on the GC injection 10 ms; info over all 54 GC injections 0.09 s. Parts are decoded at open (the largest corpus part is 278 KB); spectra parts, which can be tens of MB, are listed without being read.
Known gaps
- Spectra parts with records of tag 67 (16-bit differences) inside a
.dxare decoded as in ChemStation.uvfiles but no corpus.dxholds one; the.UVDspectra directory, MS data and any other content type are listed, not decoded. - Compound amounts, calibration, custom fields (e.g. GPC results in a peak’s
ComplexCustomFields) and system-suitability values of the.rxare not read; every corpus.rxhas unnamed compounds. - The sequence file is used for signal mapping and instrument modules only; samples, sequence lines and methods of other injections are not reported. The
.amx/.pmx/.sqxmethod files are not read. - A result-set folder is not one data set.
Vocabulary (every public identifier in openlab_*.rs must appear here)
| identifier | meaning |
|---|---|
OpenLabReader, OpenLabDataset, OPENLAB_ID, looks_like_openlab |
format reader, opened injection, the id openlab-cds, detection from the first zip member |
open, injections_in, signals, manifest, results, peak_rows, mapping, count |
open a .dx/.rx; the .dx files of a folder; decoded signals; the manifest; the results; vendor peaks with their signals; how they were mapped; values in a part |
MANIFEST_MEMBER, RESULTS_MEMBER, PAIR_BYTES, MAX_XML_BYTES |
injection.acmd; Base/InjectionACAML; 16-byte curve samples; largest XML part parsed (64 MiB) |
DxPart, name, size, compressed_size, method |
a container member: name, bytes, compressed bytes, zip method (0 stored, 8 deflate) |
DxSignal, manifest, part, part_size, header, body |
a manifest signal with its part, the part’s size and ChemStation header, the decoded body |
PartBody { Channel, Pairs, Spectra, Missing, NotDecoded }, times_ms, raw, end |
detector values; curve samples (times in ms, raw values, how the body ended); a spectra part (index of its in-memory ChemStation data set); no part; not decoded |
decode_pairs |
decode curve samples |
PeakRow, result_signal, signal, trace, peak |
a vendor peak: its result signal, signal name, our trace index, the peak |
SignalMapping { SequenceFile, BaselineMatch, SingleSignal, Unmapped } |
how result signals were matched (section above) |
InjectionManifest, version, location, injection_source, injection_volume, injection_volume_unit, sequence_line, replicate, sample_name, operator, barcode, run_started, acquisition_method |
manifest fields (table above) |
ManifestSignal, encoding, trace_id, device, device_number, channel, description, time_start, time_end, minimum, maximum, slope, declared_values, declared_records, units, integrable, kind |
manifest signal fields (declared_records: NumberOfRecords of a spectra part) |
InjectionResults, measurement_id, processing_method, processing_method_version, quantitation, software, integrator, processing_state, processed_at |
results document fields |
SignalResults, signal_id, peaks |
peaks of one result signal |
VendorPeak, peak_id, rt_min, peak_type, area, area_unit, area_percent, height, height_unit, height_percent, symmetry, start_min, end_min, baseline_code, baseline_start, baseline_end, width_base_min, baseline_model, baseline_at_apex, compound, compound_type, expected_rt_min |
one vendor peak (table above) |
SequenceFile, instrument_name, technique, modules, acquisition_software, injections, injection |
sequence file fields; find an injection |
InstrumentModule, id, manufacturer, part_number, serial_number, firmware |
one instrument module |
SequenceInjection, data_file |
one injection of the sequence file |
SequenceSignal, detector, module_id |
one signal of a sequence injection |
parse_manifest, parse_results, parse_sequence |
XML parsers |
ZipIndex, ZipMember, is_zip, first_header, norm, text, read_named, read_text, file_len, members, local_offset, encrypted, crc32 |
the shared zip reader (openreadout_core::zip) |
How this reader was derived, file by file: provenance log.