UltraIR: a foundation model for infrared spectroscopy
UltraIR learns transferable infrared (IR) spectral representations for chemical sensing and analysis, from molecules to complex samples. The project reports more than 100 million parameters and pretraining on approximately 60 million simulated IR spectra, followed by task-specific adaptation to experimental data.
This repository hosts pretrained encoder and task-adapted PyTorch .pt checkpoints, alongside data files. Use the official UltraIR Python package and the matching task configuration to load the weights.
- Code: AIMS-Lab-HKUSTGZ/UltraIR
- Paper: Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
- Authors: Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, and Jun Xia.
- Data and preprocessing: Task data guides
- License: MIT, as declared by this Hub repository and the official code license. External datasets retain their respective terms.
Model architecture and training
UltraIR uses a two-stage simulation-to-real workflow: pretrain a shared spectral encoder on simulated spectra, then jointly optimize the encoder and a task-specific head on labeled downstream data.
The encoder combines the spectrum with learned first- and second-derivative channels, gated input fusion, hierarchical convolutional feature extraction and fusion, and a patch-based Transformer. The released general pretraining configuration uses 1,792 spectral points, embedding width 1,024, patch length 16, 16 attention heads, and eight global Transformer layers. Dedicated adapters support formula-conditioned molecular structure generation and reference-guided pairwise mixture analysis.
The public pretraining implementation combines three objectives:
- Wavelet-based spectral reconstruction.
- Fingerprint-supervised representation learning using Tanimoto similarities between 2,048-bit radius-2 Morgan fingerprints.
- Prediction of 17 functional-group labels.
The released pretraining YAML specifies five epochs, training settings with learning rate 1e-4 and weight decay 0.01, and noise, shift, and masking augmentation. It provides a reproducible configuration; it does not by itself document every detail of the original large-scale run. See the pretraining implementation and paper for context.
Training data
The pretraining data guide documents UltraIR molecular-dynamics spectra, IRtoMol, the multimodal spectroscopy dataset, and QM9S, together with tools for generating molecular-dynamics spectra and combining Chemprop-IR predictions.
Prepared pretraining data consist of row-aligned ir_norm.npy ([N, L]), fingerprint.npy ([N, 2048]), and functional_groups.npy ([N, 17]). The GitHub pretraining demo contains 256 examples and is intended to exercise the pipeline.
Downstream sources include NIST, SDBS, simulated USPTO IR spectra, experimental FTIR mixtures, bacterial FTIR spectra, Jinyinhua/Shanyinhua medicinal-herb data, microplastics spectra, and the Open Soil Spectral Library. Source links, split preparation, and label definitions are in the data documentation.
Supported downstream tasks
| Area | Task | Model input and output | Configs | Reported metrics |
|---|---|---|---|---|
| Molecular interpretation | Functional-group prediction | IR spectrum to 17 multi-label functional groups | nist, sdbs, uspto |
Micro-F1, Macro-F1, exact match ratio |
| Molecular interpretation | Molecular structure elucidation | IR spectrum and molecular formula to ranked SMILES candidates | nist, sdbs, uspto |
Top-1/5/10 accuracy, validity, Tanimoto similarity, scaffold match |
| Molecular interpretation | Physicochemical property prediction | IR spectrum to 11 molecular properties | nist, sdbs, uspto |
Normalized MAE, normalized RMSE, R2 |
| Mixture analysis | Targeted component detection | Reference and mixture spectra to presence probability | nist, sdbs, uspto |
Accuracy, Macro-F1, ROC-AUC, average precision |
| Mixture analysis | Targeted fractional contribution estimation | Reference and mixture spectra to component fraction | nist, sdbs, uspto |
MAE, RMSE, R2 |
| Mixture analysis | Mixture-level component quantification | Mixture spectrum to four component quantities | experimental_four_component, synthetic_four_component |
Normalized MAE, normalized RMSE, R2 |
| Biological sensing | Bacterial classification | FTIR spectrum to one of nine genera | bacterial_classification |
Accuracy, Macro-F1, MCC |
| Botanical sensing | Medicinal-herb geographic origin traceability | FTIR spectrum to geographic origin | jyh, syh |
Accuracy, Macro-F1, MCC |
| Botanical sensing | Medicinal-herb constituent quantification | FTIR spectrum to constituent abundances | jyh_lc, syh_lc |
Normalized MAE, normalized RMSE, R2 |
| Environmental sensing | Microplastics classification | IR spectrum to one of 18 polymer classes | microplastics_classification |
Accuracy, Macro-F1, MCC |
| Environmental sensing | Soil property prediction | Mid-IR spectrum to ten soil properties | soil_property_prediction |
Normalized MAE, normalized RMSE, R2 |
The code provides NIST, SDBS, and USPTO configurations for molecular and targeted-pair tasks. The released Hub checkpoints for those tasks are NIST-specific; availability of a configuration does not imply availability of a matching fine-tuned checkpoint.
Available checkpoints
The checkpoint inventory below was checked against this Hub repository on 2026-10-02. Paths are relative to the repository root.
| Checkpoint family | Path | Files | Purpose |
|---|---|---|---|
| General encoder pretraining | checkpoints/pretraining/ultrair_pretraining_general_epoch{1..5}.pt |
5 | Initialize downstream adaptation; choose the epoch specified by the task YAML |
| Molecular-structure pretraining | checkpoints/pretraining/ultrair_pretraining_molstrelu_epoch5.pt |
1 | Initialization for the molecular-structure workflow |
| Functional-group prediction | checkpoints/functional_group_prediction/ultrair_nist.pt |
1 | NIST-adapted multi-label prediction |
| Molecular structure elucidation | checkpoints/molecular_structure_elucidation/ultrair_nist.pt |
1 | NIST-adapted formula-conditioned SMILES generation |
| Physicochemical property prediction | checkpoints/physicochemical_property_prediction/ultrair_nist.pt |
1 | NIST-adapted property regression |
| Targeted component detection | checkpoints/targeted_component_detection/ultrair_nist.pt |
1 | NIST reference/mixture presence prediction |
| Targeted fractional contribution estimation | checkpoints/targeted_fractional_contribution_estimation/ultrair_nist.pt |
1 | NIST reference/mixture fraction regression |
| Mixture-level component quantification | checkpoints/mixture_level_component_quantification/ultrair_experimental_four_component.pt |
1 | Experimental four-component mixture regression |
| Bacterial classification | checkpoints/bacterial_classification/ultrair.pt |
1 | Nine-genus classification |
| Microplastics classification | checkpoints/microplastics_classification/ultrair.pt |
1 | 18-polymer classification |
| Soil property prediction | checkpoints/soil_property_prediction/ultrair.pt |
1 | Ten-property regression |
| Medicinal-herb geographic origin | checkpoints/medicinal_herb_geographic_origin_traceability/{jyh,syh}/ultrair_{jyh,syh}_fold-{1..5}.pt |
10 | Separate herb and fold checkpoints |
| Medicinal-herb constituent quantification | checkpoints/medicinal_herb_constituent_quantification/{jyh,syh}/ultrair_{jyh,syh}_lc_fold-{1..5}.pt |
10 | Separate herb and fold checkpoints |
Brace notation denotes alternatives, not literal filenames. There are 35 checkpoints in total: six pretraining and 29 task-adapted files. The pretraining subset is approximately 3.2 GB, and the complete checkpoint tree is approximately 19.7 GB (decimal units). Downloading only the required checkpoint avoids downloading the much larger data collection.
Getting started
Install the official code
Use Python 3.11 or newer; Python 3.11 is the documented reproduction environment. The latest package installs the pinned training and data-processing dependencies from requirements.txt, including PyTorch 2.6.0, NumPy 2.2.6, RDKit 2025.3.6, PyArrow 23.0.1, and OpenCV headless 4.11.0.86. Choose a PyTorch build compatible with your CPU/CUDA environment.
The documented tested platform is Ubuntu 22.04 LTS (x86-64), Python 3.11.15, and PyTorch 2.6.0+cu124 / CUDA 12.4. The packaged demo runs on a CPU computer with 16 GB RAM. For the full USPTO pipeline, the upstream documentation specifies a CUDA-capable GPU, 16 CPU cores, and 64 GiB RAM.
git clone https://github.com/AIMS-Lab-HKUSTGZ/UltraIR.git
cd UltraIR
pip install -e .
pip install -U huggingface_hub
Core dependencies include PyTorch, NumPy, PyYAML, pytorch-wavelets, PyWavelets, RDKit, and tqdm. Run the following commands from the cloned repository root.
Download weights
Download the general epoch-5 encoder for adaptation:
hf download yusentan/UltraIR \
checkpoints/pretraining/ultrair_pretraining_general_epoch5.pt \
--local-dir .
Alternatively, download all six pretraining files with --include "checkpoints/pretraining/*.pt", or all released weights with --include "checkpoints/**". Keeping --local-dir . preserves the paths used by the YAML files.
Adapt an encoder using the packaged demo
python -m scripts.run \
--config configs/functional_group_prediction/nist.yaml \
--output-dir runs/demo \
--fold demo --epochs 1 --num-workers 0 --drop-last false \
--device cpu \
--ckpt checkpoints/pretraining/ultrair_pretraining_general_epoch5.pt
This trains and evaluates a functional-group head on the small packaged demo (70 training, 10 validation, and 20 test spectra). --drop-last false retains its 70-example training batch. Checkpoints and evaluation artifacts are saved under runs/demo/checkpoints/ and runs/demo/results/. Demo results are smoke-test outputs, not full-benchmark performance estimates. For prepared data, use --data-root /path/to/prepared/data --fold 1, or --kfold for five-fold training/evaluation. Use --output-dir runs/<experiment> to group each experiment's checkpoints and results.
Reproduce the USPTO benchmark
The new USPTO end-to-end pipeline downloads and checksum-verifies nine public IR shards, prepares 177,461 molecules and shared five-fold scaffold splits, reuses or downloads the two pretrained encoders, trains the selected tasks, reloads checkpoints for evaluation, exports prediction examples, and aggregates results.
Use the pinned Python 3.11 environment, curl, and a CUDA GPU for the full benchmark. Run either command from the repository root:
# Functional-group prediction across all five folds
python -m scripts.uspto_pipeline --work-dir runs/uspto --tasks fg
# All five USPTO molecular and targeted-mixture tasks across all five folds
python -m scripts.uspto_pipeline --work-dir runs/uspto
The full pipeline stores data under data/uspto/ and outputs under runs/uspto/. The full all-task workflow requires at least 130 GB free storage according to the upstream guide. See that guide for staged execution, task/fold selection, scratch storage, and completed-fold reuse.
USPTO here denotes the computational IR benchmark from the Zipoli, Alberts, and Laino dataset release. When using this benchmark, cite its accompanying preprint, the dataset release, and the UltraIR paper. Use RDKit 2025.03.6 and the other pinned preparation and training dependencies in requirements.txt.
Predict with a task-adapted checkpoint
For unlabeled functional-group prediction:
hf download yusentan/UltraIR \
checkpoints/functional_group_prediction/ultrair_nist.pt \
--local-dir .
python -m scripts.predict \
--config configs/functional_group_prediction/nist.yaml \
--ckpt checkpoints/functional_group_prediction/ultrair_nist.pt \
--input data/functional_group_prediction/fold-demo/test/ir_norm.npy \
--output functional_groups.json --device cpu
This outputs probabilities and selected functional groups. Functional-group prediction and targeted detection default to a threshold of 0.5; override it with --threshold. Classification outputs probabilities, regression outputs values on the original target scale, and structure elucidation outputs ranked SMILES candidates.
For a medicinal-herb model, select the statistics fold corresponding to the checkpoint:
hf download yusentan/UltraIR \
checkpoints/medicinal_herb_geographic_origin_traceability/jyh/ultrair_jyh_fold-1.pt \
--local-dir .
python -m scripts.predict \
--config configs/medicinal_herb_geographic_origin_traceability/jyh.yaml \
--ckpt checkpoints/medicinal_herb_geographic_origin_traceability/jyh/ultrair_jyh_fold-1.pt \
--stats-fold 1 \
--input /path/to/unlabeled_jyh_spectra.npy \
--output origin_predictions.json --device cpu
The packaged GitHub medicinal-herb folds supply the training reference arrays. If using a different prepared root, provide --data-root. For regression and any configuration requiring training-fold normalization, supply the matching original training reference arrays rather than demo statistics.
Input requirements
- Single-spectrum tasks accept
[L],[N, L], or[N, 1, L]NumPy inputs. - Targeted mixture tasks accept
[2, L]or[N, 2, L]: channel 0 is the pure reference, and channel 1 is the mixture. - Molecular structure elucidation additionally requires
--formula-text C5H12Oor a scalar/row-aligned formula array via--formula formula.npy. - Match the selected task's physical wavenumber range, grid direction, intensity convention, and normalization. Runtime resizing changes the point count, not the physical spectral axis.
- Molecular and generic labeled-data preparation uses row-wise min-max normalization; targeted mixtures preserve the normalized component scale without per-mixture normalization. FTIRMix quantification preserves source amplitude. Medicinal-herb preprocessing includes percent-transmission-to-absorbance conversion, per-spectrum min-max normalization, and training-fold standardization. Follow the specific data guide to avoid applying preprocessing twice.
The usual configured model input length is 1,792 points; the medicinal-herb data are resized from 1,868 points at runtime. Use the exact matching YAML rather than assuming that every task shares the same preprocessing.
For evaluation without training, use python -m scripts.evaluate --config <matching.yaml> --ckpt <task-checkpoint.pt> --fold <fold> --output-dir runs/<experiment>. --strict is appropriate when the checkpoint contains the exact complete downstream model. Full CLI instructions are in the official README.
Evaluation and limitations
The task table lists the evaluation metrics implemented by the project. Benchmark results and experimental protocols are reported in the paper. The USPTO reproduction guide documents data preparation, scaffold splits, training, checkpoint-reload evaluation, and prediction export.
This card does not assign numerical benchmark scores to individual released checkpoint files. The usage examples are instructions and were not independently rerun as part of this documentation update.
UltraIR is intended for spectroscopy research, representation learning, and adaptation to the documented analytical tasks. Generalization to new instruments, acquisition conditions, phases, sample matrices, spectral ranges, or chemical populations should be evaluated on representative labeled data. A general pretrained encoder requires a suitable task head and adaptation for new prediction targets.
Structure predictions are ranked candidates conditioned on a supplied formula; they require chemical validation. Targeted mixture prediction requires an appropriate reference spectrum, and learned fraction estimates depend on the training mixtures and signal convention. Research predictions should be validated experimentally before use in consequential analytical decisions.
Citation
@misc{tan2026simulation,
title = {Simulation-to-real transfer learning for infrared spectroscopic
chemical sensing and analysis from molecules to complex samples},
author = {Yusen Tan and Yixuan Chen and Zheng Fang and Pan Liu and
Yifan Li and Qinyu Guo and Zhedong Lin and Yuqiang Li and
Xiangxiang Zeng and Tong Wang and Jun Xia},
year = {2026},
eprint = {2608.13341},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2608.13341}
}
Sources and support
This card is based on the official GitHub README, model/pretraining code, data guides, released YAML configurations, and the Hub checkpoint inventory. It was synchronized with upstream commit 111a313 and reviewed on 2026-10-02. For questions or reproducibility issues, use GitHub Issues.
