用实验谱图训练首个纯实验型NMR结构解析模型,准确率提升17.82点。
NMRTrans: Structure Elucidation from Experimental NMR Spectra via Set Transformers
- 将谱图视为无序峰集,匹配核磁物理特性
- 在实验数据上达到61.15%的Top-10准确率,提升17.82点
- 适合需要高可靠性的化学结构解析研究者
核磁共振(NMR)光谱是分子结构解析的基础,但大规模解析仍耗时且高度依赖专家经验。现有基于谱图语言建模和检索的方法依赖大量计算谱图,对实验谱图表现显著下降。为此,我们构建了NMRSpec——一个从化学文献中挖掘的大规模$^1$H和$^{13}$C实验谱图语料库,并提出NMRTrans,将谱图建模为无序峰集,使其归纳偏置契合NMR物理本质。据我们所知,NMRTrans是首个仅使用大规模实验谱图训练的NMR Transformer,在实验基准上取得最佳性能,相比最强基线提升Top-10准确率17.82点(61.15% vs. 43.33%),凸显实验数据与结构感知架构对可靠NMR结构解析的重要性。
原文摘要 · Abstract (English)
Nuclear Magnetic Resonance (NMR) spectroscopy is fundamental for molecular structure elucidation, yet interpreting spectra at scale remains time-consuming and highly expertise-dependent. While recent spectrum-as-language modeling and retrieval-based methods have shown promise, they rely heavily on large corpora of computed spectra and exhibit notable performance drops when applied to experimental measurements. To address these issues, we build NMRSpec, a large-scale corpus of experimental $^1$H and $^{13}$C spectra mined from chemical literature, and propose NMRTrans, which models spectra as unordered peak sets and aligns the model's inductive bias with the physical nature of NMR. To our best knowledge, NMRTrans is the first NMR Transformer trained solely on large-scale experimental spectra and achieves state-of-the-art performance on experimental benchmarks, improving Top-10 Accuracy over the strongest baseline by +17.82 points (61.15% vs. 43.33%), and underscoring the importance of experimental data and structure-aware architectures for reliable NMR structure elucidation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。