用大模型统一分析多种光谱数据,自动推断分子结构
SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data
- 将不同光谱类型映射到统一语言空间,联合推理分子结构
- 在4个公开数据集上超越单模态基线,多谱联合提升准确率
- 适合需要多模态光谱分析的化学与药物研发人员
自动化分子结构解析仍具挑战,现有方法常依赖预编数据库或仅限单一光谱模态。本文提出SpectraLLM,一种可端到端预测结构的大语言模型,能基于一种或多种光谱进行推理。不同于传统谱图到结构的流程,SpectraLLM将连续(IR、Raman、UV-Vis、NMR)与离散(MS)模态统一表示于共享语言空间,捕捉跨谱型的互补子结构模式。模型在小分子领域预训练并微调,评估覆盖四个公开基准数据集。SpectraLLM实现当前最优性能,显著优于单模态基线;且在单模态设置下仍具强鲁棒性,多谱联合推理进一步提升准确率,建立可扩展的语言驱动光谱分析范式。代码已开源:https://github.com/OPilgrim/SpectraLLM。
原文摘要 · Abstract (English)
Automated molecular structure elucidation remains challenging, as existing approaches often depend on pre-compiled databases or restrict themselves to single spectroscopic modalities. Here we introduce SpectraLLM, a large language model that performs end-to-end structure prediction by reasoning over one or multiple spectra. Unlike conventional spectrum-to-structure pipelines, SpectraLLM represents both continuous (IR, Raman, UV-Vis, NMR) and discrete (MS) modalities in a shared language space, enabling it to capture substructural patterns that are complementary across different spectral types. We pretrain and fine-tune the model on small-molecule domains and evaluate it on four public benchmark datasets. SpectraLLM achieves state-of-the-art performance, substantially surpassing single-modality baselines. Moreover, it demonstrates strong robustness in unimodal settings and further improves prediction accuracy when jointly reasoning over diverse spectra, establishing a scalable paradigm for language-based spectroscopic analysis. Code is available at https://github.com/OPilgrim/SpectraLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。