arXiv:2601.18524cs.LG2026-01被引 1

用海量文献谱图无标注训练,提升核磁化学位移预测精度与泛化能力

From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale

  • 将文献中未标注的谱图作为弱监督信号,构建可扩展的半监督学习框架
  • 在数百万条谱图上实现比现有方法更优的预测准确率和跨分子类型泛化性
  • 首次大规模捕捉常见溶剂对化学位移的系统性影响,适合化学与药物研发者

精准预测核磁共振(NMR)化学位移是光谱分析与分子结构解析的基础,但现有机器学习方法依赖有限且人工标注的原子级数据集。本文提出一种半监督框架,从数百万条文献提取的未标注谱图中学习化学位移,结合少量标注数据与大规模无标签谱图。将文献谱图的化学位移预测建模为排列不变的集合监督问题,在损失函数满足常规条件下,最优二分匹配可简化为基于排序的损失,支持超越传统数据集的稳定大规模训练。模型在更广泛多样的分子数据集上显著优于当前最优方法,具备更强鲁棒性与泛化能力。此外,通过规模化引入溶剂信息,首次实现对常见NMR溶剂中系统性溶剂效应的捕捉。结果表明,从文献中挖掘的大量未标注谱图可作为训练NMR化学位移模型的有效数据源,揭示了弱结构化文献数据在科学导向数据驱动人工智能中的广阔潜力。

原文摘要 · Abstract (English)

Accurate prediction of nuclear magnetic resonance (NMR) chemical shifts is fundamental to spectral analysis and molecular structure elucidation, yet existing machine learning methods rely on limited, labor-intensive atom-assigned datasets. We propose a semi-supervised framework that learns NMR chemical shifts from millions of literature-extracted spectra without explicit atom-level assignments, integrating a small amount of labeled data with large-scale unassigned spectra. We formulate chemical shift prediction from literature spectra as a permutation-invariant set supervision problem, and show that under commonly satisfied conditions on the loss function, optimal bipartite matching reduces to a sorting-based loss, enabling stable large-scale semi-supervised training beyond traditional curated datasets. Our models achieve substantially improved accuracy and robustness over state-of-the-art methods and exhibit stronger generalization on significantly larger and more diverse molecular datasets. Moreover, by incorporating solvent information at scale, our approach captures systematic solvent effects across common NMR solvents for the first time. Overall, our results demonstrate that large-scale unlabeled spectra mined from the literature can serve as a practical and effective data source for training NMR shift models, suggesting a broader role of literature-derived, weakly structured data in data-centric AI for science.

NMR预测半监督学习文献挖掘化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。