arXiv:2502.05114cs.LGphysics.data-an2025-02

用深度模型从质谱图直接生成未知小分子结构,准确率远超传统方法。

SpecTUS: Spectral Translator for Unknown Structures annotation from EI-MS spectra

  • 直接从质谱图翻译成二维结构,无需依赖数据库
  • 单次建议正确率达43%,10次建议下达65%
  • 适合无标准库匹配的未知化合物分析

从质谱图中进行化合物识别与结构注释是一项广泛应用的任务,涵盖药物检测、刑事法医、小分子生物标志物发现及化学工程等领域。我们提出 SpecTUS:用于未知结构注释的光谱翻译器,一种深度神经模型,可针对低分辨率气相色谱-电子电离质谱(GC-EI-MS)实现小分子结构的从头注释。该方法直接将质谱图翻译为二维结构表示,特别适用于未收录于谱库的化合物分析。在多个不同谱库上的严格评估中,我们的模型显著优于传统数据库搜索方法。在包含28,267个NIST谱图的独立测试集上,单次建议可完美重建43%的化合物,且在76%的情况下优于常见的混合搜索方案。在仍具可行性的10次建议场景下,完美重建率达65%,84%的结果优于混合搜索。

原文摘要 · Abstract (English)

Compound identification and structure annotation from mass spectra is a well-established task widely applied in drug detection, criminal forensics, small molecule biomarker discovery and chemical engineering. We propose SpecTUS: Spectral Translator for Unknown Structures, a deep neural model that addresses the task of structural annotation of small molecules from low-resolution gas chromatography electron ionization mass spectra (GC-EI-MS). Our model analyzes the spectra in \textit{de novo} manner -- a direct translation from the spectra into 2D-structural representation. Our approach is particularly useful for analyzing compounds unavailable in spectral libraries. In a rigorous evaluation of our model on the novel structure annotation task across different libraries, we outperformed standard database search techniques by a wide margin. On a held-out testing set, including \numprint{28267} spectra from the NIST database, we show that our model's single suggestion perfectly reconstructs 43\% of the subset's compounds. This single suggestion is strictly better than the candidate of the database hybrid search (common method among practitioners) in 76\% of cases. In a~still affordable scenario of~10 suggestions, perfect reconstruction is achieved in 65\%, and 84\% are better than the hybrid search.

质谱分析结构预测深度学习从头注释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。