JCAMP-DX
JCAMP-DX is the open IUPAC text format for spectra, exported by many instruments and programs for IR, NMR, mass and other spectra. OpenReadout returns each data table as a trace with its x axis and the descriptive records, and export --to jcamp writes NMR and other 1-D spectra as JCAMP-DX. Derived from the public IUPAC JCAMP-DX protocols (IR 4.24, 1988; NMR, 1993; MS, 1994; 5.01, 1999) and checked on public test files (ISAS Dortmund test suite, R. J. Lancashire’s public-domain files, MestReNova 14 exports); nmrglue’s jcampdx (BSD-3) and the jcamp package (MIT) were read as prior art and serve as reference readers. Provenance: docs/provenance/jcamp-dx.md.
Format id jcamp-dx, family spectroscopy; extensions jdx, dx, jcamp, jcm. Section numbers (§) refer to the 4.24 protocol unless stated.
Structure (parse_jcamp, JcampFile, Block, Ldr)
- Printable text in labeled data records (LDRs):
##LABEL= value, continued on the following lines until the next##(§4.2). Labels are compared after upper-casing and dropping spaces, dashes, slashes and underscores (§4.4,normalize_label:DATA TYPE,DATATYPEandData_Typeare one label).$$starts a comment (removed from values). User-defined labels start with$(##$AQ_mod=in Bruker exports); data-type-specific labels start with.(##.OBSERVE FREQUENCY=). - A block runs from
##TITLE=to its##END=. Blocks nest: a##DATA TYPE= LINKblock of a compound file holds child blocks (§6.1.4;parent,depth). Unbalanced##END=or a block without one are reported; a block without##END=is atruncatederror. A bare##END(no=) closes the block too (end_without_equals, info). - Lines may end in LF, CRLF or a lone CR (classic Mac,
jcamp-lancashire-mactab2); blanks before##are accepted (TESTNTUP.DX). Text that is not UTF-8 is read as Latin-1 (latin1). - Detection: the file starts (after blanks and a byte-order mark) with
##TITLE=→ definite; only the extension → extension-only.
Data tables and how they become traces
One trace per data table, in file order (LINK blocks have none). Values are the table values multiplied by the declared factor; the factor is the channel scale.
| table | variable list | trace |
|---|---|---|
##XYDATA= (§6.4.1) |
(X++(Y..Y)), (X++(R..R)), (X++(I..I)) (NMR §4.1.1) |
one channel (y, real, imag), factor ##YFACTOR=; several ##XYDATA= in one block become channels; X = ##FIRSTX= + i·(##LASTX=−##FIRSTX=)/(##NPOINTS=−1) in extra.axis |
##NTUPLES= (NMR §7, 5.01) |
per page ##DATA TABLE= (X++(R..R)), XYDATA |
pages keyed N=1, N=2 holding different symbols (R, I) become channels; pages keyed by another variable (F1=17094.6, a retention time) become sweeps (extra.sweep_axis); channel names are the ##VAR_NAME= entries, factors ##FACTOR=, units ##UNITS=; the independent variable’s ##FIRST=/##LAST=/##VAR_DIM= give the axis |
##PEAK TABLE=, ##XYPOINTS= (§6.4.2–3) |
(XY..XY), (XYW..XYW), (XYM..XYM) |
channels x, y [, w/m] (the abscissa is a channel because it is not evenly spaced; extra.axis.irregular = true); factors ##XFACTOR=/##YFACTOR=; non-numeric components (multiplicity letters, <strings>) are NaN |
NTUPLES with (XY..XY) pages (MS spectral series) |
channels from the group symbols, one sweep per page; pages may differ in length (sample_count is the largest; extra.sweep_sample_counts lists each page’s count when they differ) |
|
##PEAK ASSIGNMENTS=, ##RADATA=, JCAMP-CS structure blocks |
listed by info --view structure, no trace |
sample_count: ##NPOINTS= (XYDATA, peak tables) or the independent variable’s ##VAR_DIM= (NTUPLES). sample_rate_hz: (n−1)/|LASTX−FIRSTX| when the X unit is time (SECONDS, MS, MINUTES), else 0. start_s: FIRSTX (or the NTUPLES ##FIRST=) in seconds when the X unit is time and the abscissa increases, else unset (a falling time axis is placed by extra.axis only).
ASDF (decode_asdf, lex_asdf_line; §5)
Each line: an abscissa (AFFN), then ordinates. Tokens (YToken):
| form | characters | meaning |
|---|---|---|
| AFFN / PAC | +, -, digits, ., exponent E±dd |
actual value (Abs); a sign also separates values (PAC) |
| SQZ | @ A–I = 0…9, a–i = −1…−9 |
actual value, leading digit and sign in one character (Abs) |
| DIF | % J–R = 0…9, j–r = −1…−9 |
difference from the previous ordinate (Dif) |
| DUP | S–Z = 1…8, s = 9 |
the previous token occurs this many times in total (Dup; in DIFDUP the difference repeats) |
? |
missing or off-scale ordinate (Invalid, NaN) |
E/e is an exponent only when followed by a sign and a digit; otherwise it is SQZ 5 / −5.
Check-points (§5.8). When a line ends in DIF form (including a DUP of a DIF), the next line’s first ordinate repeats the last ordinate (Y-value check): it is compared and dropped, and a DUP right after it repeats that value (TESTNTUP.DX writes h5T). A mismatch is a y_check error. Each line’s abscissa labels the ordinate at its index (the repeated one after a DIF line); abscissas more than half a step plus their own rounding (half a unit in their last written digit, times XFACTOR) from FIRSTX + i·step are an x_sequence warning (MestReNova writes whole Hz; Spectrafile-IR labels the first new point). The last line of a DIF table may hold only the check value.
Descriptive records copied into trace extra (our name ← label)
| our name | label | notes |
|---|---|---|
kind |
– | xydata, ntuples, peak_table, xypoints |
block |
– | block index (see info --view structure) |
title |
TITLE |
also the trace name |
jcamp_version |
JCAMP-DX |
|
data_type, data_class |
DATA TYPE, DATA CLASS |
|
origin, owner |
ORIGIN, OWNER |
|
spectrometer |
SPECTROMETER/DATA SYSTEM |
|
instrument_parameters |
INSTRUMENT PARAMETERS |
|
sample_description |
SAMPLE DESCRIPTION |
|
compound_name, names, molecular_formula, cas_registry_number |
CAS NAME, NAMES, MOLFORM, CAS REGISTRY NO |
|
date, time, long_date |
DATE, TIME, LONG DATE |
verbatim |
nucleus |
.OBSERVE NUCLEUS |
e.g. ^13C |
observe_frequency_mhz |
.OBSERVE FREQUENCY |
|
field_t |
.FIELD |
|
scans |
.AVERAGES |
|
solvent |
.SOLVENT NAME |
|
pulse_sequence |
.PULSE SEQUENCE |
|
acquisition_mode |
.ACQUISITION MODE |
|
ms_spectrometer_type, ms_inlet, ms_ionization_mode |
.SPECTROMETER TYPE, .INLET, .IONIZATION MODE |
MS protocol §5.2 |
resolution, min_y, max_y, first_y |
RESOLUTION, MINY, MAXY, FIRSTY |
numbers |
axis |
XUNITS/FIRSTX/LASTX/NPOINTS or NTUPLES UNITS/FIRST/LAST/VAR_DIM |
{quantity, unit, first, last, step, size}; quantity from the unit: frequency (HZ), chemical_shift (PPM), time, wavenumber (1/CM), wavelength (NANOMETERS, MICROMETERS), mass_to_charge (M/Z), else x |
sweep_axis |
NTUPLES PAGE= |
{variable, unit, values} |
sweep_sample_counts |
– | points per page when NTUPLES pages differ in length |
ntuples, symbols, variables |
NTUPLES, SYMBOL, VAR_NAME |
|
undecodable, problems, note |
– | why a table cannot be read |
Records of the enclosing LINK block apply to its children (a compound file’s shared ##ORIGIN=), except TITLE, DATA TYPE and DATA CLASS. Every record is in info --view full’s vendor tree (blocks[].records, tables summarized as their variable list and line count).
check finding codes
truncated, no_blocks, y_check, npoints_mismatch, bad_table, missing_label (XYDATA without FIRSTX/LASTX/NPOINTS) (errors); x_sequence, firsty_mismatch (tolerance: one factor step), missing_label, outside_block, extra_end, bad_label, undecodable, no_data (warnings); non_numeric, non_utf8, end_without_equals (info).
Observed corpus values
| id | writer | table | notes |
|---|---|---|---|
jcamp-isas-brukaffn, -bruksqz, -brukpac, -brukdif |
Bruker NMR JCAMP-DX V1.0 (1992) | XYDATA 16384, AFFN / SQZ / PAC / DIFDUP | AFFN, SQZ and PAC decode to identical values |
jcamp-isas-brukntup, -testntup |
Bruker, ISAS | NTUPLES R/I, 16384 | TESTNTUP: X column is an index × ##FACTOR= |
jcamp-isas-testfid |
ISAS | NTUPLES FID (TIME, FID/REAL, FID/IMAG) | |
jcamp-isas-pe1800, -specfile |
Perkin Elmer 1800, Spectrafile-IR | IR XYDATA PAC, DIFDUP | SPECFILE: S (DUP 1) after every value; last line 31999@ fails its check (0 ≠ 26506): check exit 4 |
jcamp-isas-ms1, -ms2 |
ISAS MS | PEAK TABLE; XYDATA with descending X in seconds | ##DATA CLASS= PEAKTABLE |
jcamp-lancashire-compound, -blckpac1 |
UWI Mona | LINK with 5 IR / UV-Vis XYDATA blocks | |
jcamp-lancashire-dupinc1, -pacdec1, -sqzdec1 |
Perkin Elmer, T. Davies | UV-Vis DUP, IR PAC, NMR SQZ | |
jcamp-lancashire-pktab1, -mactab2 |
UWI Mona | MS PEAK TABLE | mactab2: CR line ends and a trailing 0xFF byte |
nmrxiv-s200-qhnmr |
MestReNova 14.0.1 (JCAMP-DX 6.0) | LINK + XYDATA 104858 points | X written in whole Hz |
nmrxiv-s200-hsqc |
MestReNova 14.0.1 | nD NMR SPECTRUM NTUPLES, 1024 pages F1= × 820 points (F2++(Y..Y)) |
##FIRST= repeated in every page |
Writing (export --to jcamp; export_jcamp, jcamp_write.rs)
One trace per file, JCAMP-DX 5.01 (the NMR protocol’s layout, readable by the IR/MS protocol readers):
| trace | written as |
|---|---|
| one channel, one sweep, regular axis | ##DATA CLASS= XYDATA, ##XYDATA= (X++(Y..Y)) with XUNITS/YUNITS, XFACTOR, YFACTOR, FIRSTX, LASTX, DELTAX, NPOINTS, FIRSTY |
| several channels or sweeps | ##DATA CLASS= NTUPLES: ##NTUPLES=, VAR_NAME, SYMBOL (X, then R/I for a real/imag pair, Y for one channel, Y1, Y2, … otherwise, and N), VAR_TYPE, VAR_FORM, VAR_DIM, UNITS, FIRST, LAST, FACTOR; one ##PAGE= N=k + ##DATA TABLE= (X++(R..R)), XYDATA per channel and sweep (one sweep: N numbers the channels, as NMR FIDs do; several sweeps: N numbers the sweeps and every sweep has one page per channel) |
irregular abscissa (extra.axis.irregular: peak tables, point lists), two columns, one sweep |
##PEAK TABLE= (XY..XY) (or ##XYPOINTS=) with AFFN x,y pairs in their shortest exact decimal form, factors 1 |
Header: TITLE (trace name), JCAMP-DX= 5.01, DATA TYPE (the source’s own for JCAMP-DX input; NMR FID/NMR SPECTRUM for Bruker time-domain/processed traces; CHROMATOGRAM for chromatography; UNKNOWN otherwise), DATA CLASS, ORIGIN, OWNER (source or experiment model; (not recorded) when unknown), LONGDATE (YYYY/MM/DD HH:MM:SS from the acquisition start), SPECTROMETER/DATA SYSTEM, .OBSERVE FREQUENCY, .OBSERVE NUCLEUS (^1H), .SOLVENT NAME, .PULSE SEQUENCE, and ##$OPENREADOUT SOURCE= (format id, file name, trace).
Ordinates are integers × a per-channel factor (YFACTOR or the NTUPLES FACTOR). The factor is, in order: the channel’s own scale (a JCAMP-DX input’s YFACTOR) when every value is exactly an integer multiple of it; else the largest power of two that makes every value an integer (Bruker values are integers × 2^NC), provided the integers stay within 2^53; else (values spanning more than 53 bits, e.g. decimal fractions) a power of two that fits the largest magnitude in 32 bits, and the report says exact: false with max_abs_error ≤ factor / 2. Non-finite values are refused (exit 6).
DIFDUP (default): each line is the abscissa as an integer multiple of XFACTOR = |step| (the Bruker convention), the first ordinate in SQZ form and differences in DIF form, repeats of a difference as DUP counts of one character (≤ 9 per group, because the jcamp Python package reads single-character counts only); lines are at most 80 characters and each ends in DIF form, so the next line starts with the Y-value check (the last ordinate again); a final line repeats the last ordinate as the last check. --compression none writes AFFN instead (space-separated integers).
Verification: the file is re-read with JcampDataset: one trace of the written shape, every ordinate equal to integer × factor bit for bit, the axis within 1e-9 relative, and check without errors; only then is it renamed into place. Third-party readers: oracle/jcamp_validate.py (nmrglue for NMR data types, jcamp for XYDATA and peak tables; jcamp cannot read negative numbers in these tables, so files whose abscissa goes below zero, e.g. most ppm spectra, are read by nmrglue only).
Vocabulary (every public identifier in crates/openreadout-nmr/src/jcamp_*.rs must appear here)
| identifier | meaning |
|---|---|
JcampReader, JcampDataset, JCAMP_FORMAT_ID, open |
reader entry points: format reader, opened file (core Dataset), the id jcamp-dx |
JcampFile, blocks, issues, latin1, len, parse_jcamp |
a parsed file: blocks, structural findings, Latin-1 fallback, size |
Block, index, parent, depth, ldrs, offset, end, closed |
a ##TITLE= … ##END= block |
get, all, text, number, title, data_type, data_class, is_link |
record lookup and common labels |
Ldr, label, key, value, line, head, body |
one labeled data record: label as written, normalized key, value, line number; first line and table body |
normalize_label, parse_affn |
label normalization (§4.4) and AFFN numbers (§4.5.3) |
YToken { Abs, Dif, Dup, Invalid } |
ASDF ordinate tokens |
AsdfLine, x, x_resolution, tokens, lex_asdf_line |
one lexed data line |
AsdfTable, y, x_checks, y_check_failures, lines, decode_asdf |
a decoded (X++(Y..Y)) table and its check-points |
decode_groups |
(XY..XY) group tables |
export_jcamp, default_jcamp_output |
write one trace as JCAMP-DX (verified by re-reading); the default output name <stem>[.traceT][.sweepS].jdx |
JcampExportOptions, rows, encoding, overwrite |
what to write: trace, sweep, samples [first, last] of each sweep, DIFDUP or AFFN, replace an existing file |
JcampEncoding { Difdup, Affn } |
ordinate form |
JcampExportReport, input, output, jcamp_version, data_class, sweeps, pages, first_sample, samples_written, channels_written, factors, exact, max_abs_error, bytes_written, verified |
the export report (format, data_type, trace as elsewhere in this table) |
How this reader was derived, file by file: provenance log.