用多模态光谱数据自动推断有机分子结构,准确率达93.8%。
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

- 构建光谱-分子联合学习框架,融合多模态谱图与大规模分子先验。
- 在模拟数据上实现93.8%的单模型预测准确率,实验数据经微调后进一步提升。
- 适合化学信息学、药物发现及自动化结构解析研究者使用。
从光谱数据中确定分子结构仍具根本挑战,因逆问题本质欠定:单个谱图稀疏、低维,仅提供相对于庞大分子空间的部分结构证据。我们提出一种可扩展的假设-精炼范式,将光谱证据与大规模分子先验紧密结合。为提供结构分辨的NMR信号,构建了基于DFT的QM9SPIN数据集,包含多样化的1D和2D谱图,如J耦合、DEPT实验及显式自旋-自旋相互作用。在此基础上,提出SpectroMol模型,根据多模态光谱输入生成化学有效分子假设。同时开发高分辨率质量约束生成器MS-Mol2Mol,结合分子式、精确质量与不饱和度,在4亿分子训练的条件生成先验中确保全局组成一致性与化学合理性。集成系统在模拟基准测试中达到93.8%的top-1准确率,能有效从模拟谱图迁移至实验谱图,且通过质量引导精炼进一步提升实验预测性能,为自动化、数据驱动的有机结构解析提供可扩展路径。
原文摘要 · Abstract (English)
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。