用1.58亿组模拟NMR数据训练大模型,实现从仿真到真实谱图的精准分析
A large-scale foundation model enables simulation-to-real adaptation for nuclear magnetic resonance-based molecular structure analysis

- 基于1.58亿对模拟氢碳核磁数据,构建可泛化的谱图表征模型
- 在真实样品分析中性能超越现有方法,成功解析两种未知中药成分
- 适合需要快速结构鉴定的药物研发与天然产物研究者使用
核磁共振(NMR)光谱是分子结构分析的强大工具,而光谱人工智能在快速自动化解析方面潜力巨大。然而,实验NMR数据集稀缺限制了深度学习在此领域的应用,仅限于特定任务且泛化能力弱。本文提出UltraNMR,一种大规模基础模型,利用NMR谱图的内在特性学习通用谱图表示。我们收集了1.58亿对模拟¹H和¹³C NMR谱图用于训练,采用多种领域特定预训练目标。UltraNMR捕捉谱图内与谱图间依赖关系,实现无缝仿真到真实的适应。实验证明,将UltraNMR适配至多种真实分子结构分析任务时,性能持续达到顶尖水平,并显著优于直接在下游数据上训练的变体。我们还通过UltraNMR编码模拟谱图,构建了包含9400万唯一分子的大规模谱图向量库,支持结构感知检索。在实际应用中,UltraNMR成功协助解析两种来自《中国药典》记载的未知中药成分。结果表明,大规模模拟预训练能有效弥合仿真与真实之间的差距,实现稳健且通用的现实世界NMR谱图分析。
原文摘要 · Abstract (English)
Nuclear Magnetic Resonance (NMR) spectroscopy is a powerful tool for molecular structure analysis, and spectral artificial intelligence offers great potential for its rapid and automated interpretation. However, the scarcity of experimental NMR datasets has constrained deep learning in this domain to narrow, task-specific applications that lack broad generalization. Here, we introduce UltraNMR, a large-scale foundation model for NMR that leverages the intrinsic properties of NMR spectra to learn generalizable spectral representations. We collected 158 million paired simulated $^{1}$H and $^{13}$C NMR spectra to train UltraNMR, employing multiple domain-specific pre-training objectives. UltraNMR captures both intra-spectral and inter-spectral dependencies, enabling seamless simulation-to-real adaptation. We demonstrate that adapting UltraNMR to a range of molecular structure analysis tasks on experimental NMR spectra consistently yields state-of-the-art performance and clearly outperforms UltraNMR variants trained directly on downstream data without simulation pre-training. We also construct a large-scale NMR spectral vector library by encoding simulated NMR spectra using UltraNMR, covering 94 million unique molecules and enabling effective structure-aware retrieval. In real-world applications, UltraNMR facilitates the structural elucidation of two previously unknown natural products from Chinese herbal medicines recorded in the Chinese Pharmacopoeia. These results suggest that large-scale simulation pre-training can effectively bridge the simulation-to-real gap, enabling robust and generalizable molecular structure analysis of real-world NMR spectra.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。