UltraIR: a foundation model for infrared spectroscopy

UltraIR learns transferable infrared (IR) spectral representations for chemical sensing and analysis, from molecules to complex samples. The project reports more than 100 million parameters and pretraining on approximately 60 million simulated IR spectra, followed by task-specific adaptation to experimental data.

This repository hosts pretrained encoder and task-adapted PyTorch .pt checkpoints, alongside data files. Use the official UltraIR Python package and the matching task configuration to load the weights.

UltraIR framework and encoder architecture

Model architecture and training

UltraIR uses a two-stage simulation-to-real workflow: pretrain a shared spectral encoder on simulated spectra, then jointly optimize the encoder and a task-specific head on labeled downstream data.

The encoder combines the spectrum with learned first- and second-derivative channels, gated input fusion, hierarchical convolutional feature extraction and fusion, and a patch-based Transformer. The released general pretraining configuration uses 1,792 spectral points, embedding width 1,024, patch length 16, 16 attention heads, and eight global Transformer layers. Dedicated adapters support formula-conditioned molecular structure generation and reference-guided pairwise mixture analysis.

The public pretraining implementation combines three objectives:

  1. Wavelet-based spectral reconstruction.
  2. Fingerprint-supervised representation learning using Tanimoto similarities between 2,048-bit radius-2 Morgan fingerprints.
  3. Prediction of 17 functional-group labels.

The released pretraining YAML specifies five epochs, training settings with learning rate 1e-4 and weight decay 0.01, and noise, shift, and masking augmentation. It provides a reproducible configuration; it does not by itself document every detail of the original large-scale run. See the pretraining implementation and paper for context.

Training data

The pretraining data guide documents UltraIR molecular-dynamics spectra, IRtoMol, the multimodal spectroscopy dataset, and QM9S, together with tools for generating molecular-dynamics spectra and combining Chemprop-IR predictions.

Prepared pretraining data consist of row-aligned ir_norm.npy ([N, L]), fingerprint.npy ([N, 2048]), and functional_groups.npy ([N, 17]). The GitHub pretraining demo contains 256 examples and is intended to exercise the pipeline.

Downstream sources include NIST, SDBS, simulated USPTO IR spectra, experimental FTIR mixtures, bacterial FTIR spectra, Jinyinhua/Shanyinhua medicinal-herb data, microplastics spectra, and the Open Soil Spectral Library. Source links, split preparation, and label definitions are in the data documentation.

Supported downstream tasks

Area Task Model input and output Configs Reported metrics
Molecular interpretation Functional-group prediction IR spectrum to 17 multi-label functional groups nist, sdbs, uspto Micro-F1, Macro-F1, exact match ratio
Molecular interpretation Molecular structure elucidation IR spectrum and molecular formula to ranked SMILES candidates nist, sdbs, uspto Top-1/5/10 accuracy, validity, Tanimoto similarity, scaffold match
Molecular interpretation Physicochemical property prediction IR spectrum to 11 molecular properties nist, sdbs, uspto Normalized MAE, normalized RMSE, R2
Mixture analysis Targeted component detection Reference and mixture spectra to presence probability nist, sdbs, uspto Accuracy, Macro-F1, ROC-AUC, average precision
Mixture analysis Targeted fractional contribution estimation Reference and mixture spectra to component fraction nist, sdbs, uspto MAE, RMSE, R2
Mixture analysis Mixture-level component quantification Mixture spectrum to four component quantities experimental_four_component, synthetic_four_component Normalized MAE, normalized RMSE, R2
Biological sensing Bacterial classification FTIR spectrum to one of nine genera bacterial_classification Accuracy, Macro-F1, MCC
Botanical sensing Medicinal-herb geographic origin traceability FTIR spectrum to geographic origin jyh, syh Accuracy, Macro-F1, MCC
Botanical sensing Medicinal-herb constituent quantification FTIR spectrum to constituent abundances jyh_lc, syh_lc Normalized MAE, normalized RMSE, R2
Environmental sensing Microplastics classification IR spectrum to one of 18 polymer classes microplastics_classification Accuracy, Macro-F1, MCC
Environmental sensing Soil property prediction Mid-IR spectrum to ten soil properties soil_property_prediction Normalized MAE, normalized RMSE, R2

The code provides NIST, SDBS, and USPTO configurations for molecular and targeted-pair tasks. The released Hub checkpoints for those tasks are NIST-specific; availability of a configuration does not imply availability of a matching fine-tuned checkpoint.

Available checkpoints

The checkpoint inventory below was checked against this Hub repository on 2026-10-02. Paths are relative to the repository root.

Checkpoint family Path Files Purpose
General encoder pretraining checkpoints/pretraining/ultrair_pretraining_general_epoch{1..5}.pt 5 Initialize downstream adaptation; choose the epoch specified by the task YAML
Molecular-structure pretraining checkpoints/pretraining/ultrair_pretraining_molstrelu_epoch5.pt 1 Initialization for the molecular-structure workflow
Functional-group prediction checkpoints/functional_group_prediction/ultrair_nist.pt 1 NIST-adapted multi-label prediction
Molecular structure elucidation checkpoints/molecular_structure_elucidation/ultrair_nist.pt 1 NIST-adapted formula-conditioned SMILES generation
Physicochemical property prediction checkpoints/physicochemical_property_prediction/ultrair_nist.pt 1 NIST-adapted property regression
Targeted component detection checkpoints/targeted_component_detection/ultrair_nist.pt 1 NIST reference/mixture presence prediction
Targeted fractional contribution estimation checkpoints/targeted_fractional_contribution_estimation/ultrair_nist.pt 1 NIST reference/mixture fraction regression
Mixture-level component quantification checkpoints/mixture_level_component_quantification/ultrair_experimental_four_component.pt 1 Experimental four-component mixture regression
Bacterial classification checkpoints/bacterial_classification/ultrair.pt 1 Nine-genus classification
Microplastics classification checkpoints/microplastics_classification/ultrair.pt 1 18-polymer classification
Soil property prediction checkpoints/soil_property_prediction/ultrair.pt 1 Ten-property regression
Medicinal-herb geographic origin checkpoints/medicinal_herb_geographic_origin_traceability/{jyh,syh}/ultrair_{jyh,syh}_fold-{1..5}.pt 10 Separate herb and fold checkpoints
Medicinal-herb constituent quantification checkpoints/medicinal_herb_constituent_quantification/{jyh,syh}/ultrair_{jyh,syh}_lc_fold-{1..5}.pt 10 Separate herb and fold checkpoints

Brace notation denotes alternatives, not literal filenames. There are 35 checkpoints in total: six pretraining and 29 task-adapted files. The pretraining subset is approximately 3.2 GB, and the complete checkpoint tree is approximately 19.7 GB (decimal units). Downloading only the required checkpoint avoids downloading the much larger data collection.

Getting started

Install the official code

Use Python 3.11 or newer; Python 3.11 is the documented reproduction environment. The latest package installs the pinned training and data-processing dependencies from requirements.txt, including PyTorch 2.6.0, NumPy 2.2.6, RDKit 2025.3.6, PyArrow 23.0.1, and OpenCV headless 4.11.0.86. Choose a PyTorch build compatible with your CPU/CUDA environment.

The documented tested platform is Ubuntu 22.04 LTS (x86-64), Python 3.11.15, and PyTorch 2.6.0+cu124 / CUDA 12.4. The packaged demo runs on a CPU computer with 16 GB RAM. For the full USPTO pipeline, the upstream documentation specifies a CUDA-capable GPU, 16 CPU cores, and 64 GiB RAM.

git clone https://github.com/AIMS-Lab-HKUSTGZ/UltraIR.git
cd UltraIR
pip install -e .
pip install -U huggingface_hub

Core dependencies include PyTorch, NumPy, PyYAML, pytorch-wavelets, PyWavelets, RDKit, and tqdm. Run the following commands from the cloned repository root.

Download weights

Download the general epoch-5 encoder for adaptation:

hf download yusentan/UltraIR \
  checkpoints/pretraining/ultrair_pretraining_general_epoch5.pt \
  --local-dir .

Alternatively, download all six pretraining files with --include "checkpoints/pretraining/*.pt", or all released weights with --include "checkpoints/**". Keeping --local-dir . preserves the paths used by the YAML files.

Adapt an encoder using the packaged demo

python -m scripts.run \
  --config configs/functional_group_prediction/nist.yaml \
  --output-dir runs/demo \
  --fold demo --epochs 1 --num-workers 0 --drop-last false \
  --device cpu \
  --ckpt checkpoints/pretraining/ultrair_pretraining_general_epoch5.pt

This trains and evaluates a functional-group head on the small packaged demo (70 training, 10 validation, and 20 test spectra). --drop-last false retains its 70-example training batch. Checkpoints and evaluation artifacts are saved under runs/demo/checkpoints/ and runs/demo/results/. Demo results are smoke-test outputs, not full-benchmark performance estimates. For prepared data, use --data-root /path/to/prepared/data --fold 1, or --kfold for five-fold training/evaluation. Use --output-dir runs/<experiment> to group each experiment's checkpoints and results.

Reproduce the USPTO benchmark

The new USPTO end-to-end pipeline downloads and checksum-verifies nine public IR shards, prepares 177,461 molecules and shared five-fold scaffold splits, reuses or downloads the two pretrained encoders, trains the selected tasks, reloads checkpoints for evaluation, exports prediction examples, and aggregates results.

Use the pinned Python 3.11 environment, curl, and a CUDA GPU for the full benchmark. Run either command from the repository root:

# Functional-group prediction across all five folds
python -m scripts.uspto_pipeline --work-dir runs/uspto --tasks fg

# All five USPTO molecular and targeted-mixture tasks across all five folds
python -m scripts.uspto_pipeline --work-dir runs/uspto

The full pipeline stores data under data/uspto/ and outputs under runs/uspto/. The full all-task workflow requires at least 130 GB free storage according to the upstream guide. See that guide for staged execution, task/fold selection, scratch storage, and completed-fold reuse.

USPTO here denotes the computational IR benchmark from the Zipoli, Alberts, and Laino dataset release. When using this benchmark, cite its accompanying preprint, the dataset release, and the UltraIR paper. Use RDKit 2025.03.6 and the other pinned preparation and training dependencies in requirements.txt.

Predict with a task-adapted checkpoint

For unlabeled functional-group prediction:

hf download yusentan/UltraIR \
  checkpoints/functional_group_prediction/ultrair_nist.pt \
  --local-dir .

python -m scripts.predict \
  --config configs/functional_group_prediction/nist.yaml \
  --ckpt checkpoints/functional_group_prediction/ultrair_nist.pt \
  --input data/functional_group_prediction/fold-demo/test/ir_norm.npy \
  --output functional_groups.json --device cpu

This outputs probabilities and selected functional groups. Functional-group prediction and targeted detection default to a threshold of 0.5; override it with --threshold. Classification outputs probabilities, regression outputs values on the original target scale, and structure elucidation outputs ranked SMILES candidates.

For a medicinal-herb model, select the statistics fold corresponding to the checkpoint:

hf download yusentan/UltraIR \
  checkpoints/medicinal_herb_geographic_origin_traceability/jyh/ultrair_jyh_fold-1.pt \
  --local-dir .

python -m scripts.predict \
  --config configs/medicinal_herb_geographic_origin_traceability/jyh.yaml \
  --ckpt checkpoints/medicinal_herb_geographic_origin_traceability/jyh/ultrair_jyh_fold-1.pt \
  --stats-fold 1 \
  --input /path/to/unlabeled_jyh_spectra.npy \
  --output origin_predictions.json --device cpu

The packaged GitHub medicinal-herb folds supply the training reference arrays. If using a different prepared root, provide --data-root. For regression and any configuration requiring training-fold normalization, supply the matching original training reference arrays rather than demo statistics.

Input requirements

  • Single-spectrum tasks accept [L], [N, L], or [N, 1, L] NumPy inputs.
  • Targeted mixture tasks accept [2, L] or [N, 2, L]: channel 0 is the pure reference, and channel 1 is the mixture.
  • Molecular structure elucidation additionally requires --formula-text C5H12O or a scalar/row-aligned formula array via --formula formula.npy.
  • Match the selected task's physical wavenumber range, grid direction, intensity convention, and normalization. Runtime resizing changes the point count, not the physical spectral axis.
  • Molecular and generic labeled-data preparation uses row-wise min-max normalization; targeted mixtures preserve the normalized component scale without per-mixture normalization. FTIRMix quantification preserves source amplitude. Medicinal-herb preprocessing includes percent-transmission-to-absorbance conversion, per-spectrum min-max normalization, and training-fold standardization. Follow the specific data guide to avoid applying preprocessing twice.

The usual configured model input length is 1,792 points; the medicinal-herb data are resized from 1,868 points at runtime. Use the exact matching YAML rather than assuming that every task shares the same preprocessing.

For evaluation without training, use python -m scripts.evaluate --config <matching.yaml> --ckpt <task-checkpoint.pt> --fold <fold> --output-dir runs/<experiment>. --strict is appropriate when the checkpoint contains the exact complete downstream model. Full CLI instructions are in the official README.

Evaluation and limitations

The task table lists the evaluation metrics implemented by the project. Benchmark results and experimental protocols are reported in the paper. The USPTO reproduction guide documents data preparation, scaffold splits, training, checkpoint-reload evaluation, and prediction export.

This card does not assign numerical benchmark scores to individual released checkpoint files. The usage examples are instructions and were not independently rerun as part of this documentation update.

UltraIR is intended for spectroscopy research, representation learning, and adaptation to the documented analytical tasks. Generalization to new instruments, acquisition conditions, phases, sample matrices, spectral ranges, or chemical populations should be evaluated on representative labeled data. A general pretrained encoder requires a suitable task head and adaptation for new prediction targets.

Structure predictions are ranked candidates conditioned on a supplied formula; they require chemical validation. Targeted mixture prediction requires an appropriate reference spectrum, and learned fraction estimates depend on the training mixtures and signal convention. Research predictions should be validated experimentally before use in consequential analytical decisions.

Citation

@misc{tan2026simulation,
  title         = {Simulation-to-real transfer learning for infrared spectroscopic
                   chemical sensing and analysis from molecules to complex samples},
  author        = {Yusen Tan and Yixuan Chen and Zheng Fang and Pan Liu and
                   Yifan Li and Qinyu Guo and Zhedong Lin and Yuqiang Li and
                   Xiangxiang Zeng and Tong Wang and Jun Xia},
  year          = {2026},
  eprint        = {2608.13341},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2608.13341}
}

Sources and support

This card is based on the official GitHub README, model/pretraining code, data guides, released YAML configurations, and the Hub checkpoint inventory. It was synchronized with upstream commit 111a313 and reviewed on 2026-10-02. For questions or reproducibility issues, use GitHub Issues.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for yusentan/UltraIR