arXiv:2602.16507cs.LG2026-02被引 6

深度学习预测质谱分子指纹时,准确度与检索效果存在根本矛盾。

Small molecule retrieval from tandem mass spectrometry: what are we optimizing for?

  • 用损失函数优化分子指纹预测
  • 指纹越准,化合物检索反而越差
  • 理论揭示指纹相似与检索间的权衡

液相色谱-串联质谱(LC-MS/MS)数据计算分析的核心挑战之一是识别输出谱图对应的化合物。近年来,该问题越来越多地采用深度学习方法解决。常见策略是从输入质谱预测分子指纹向量,并在化学数据库中搜索匹配项。尽管训练中使用多种损失函数,其对模型性能的影响仍不明确。本研究分析了常用损失函数,推导出新的后悔界,揭示了贝叶斯最优决策在不同目标下可能分歧。结果表明,指纹相似性与分子检索之间存在根本性权衡:优化指纹预测精度通常会损害检索效果,反之亦然。理论分析显示,这一权衡取决于候选集的相似性结构,为损失函数和指纹选择提供指导。

原文摘要 · Abstract (English)

One of the central challenges in the computational analysis of liquid chromatography-tandem mass spectrometry (LC-MS/MS) data is to identify the compounds underlying the output spectra. In recent years, this problem is increasingly tackled using deep learning methods. A common strategy involves predicting a molecular fingerprint vector from an input mass spectrum, which is then used to search for matches in a chemical compound database. While various loss functions are employed in training these predictive models, their impact on model performance remains poorly understood. In this study, we investigate commonly used loss functions, deriving novel regret bounds that characterize when Bayes-optimal decisions for these objectives must diverge. Our results reveal a fundamental trade-off between the two objectives of (1) fingerprint similarity and (2) molecular retrieval. Optimizing for more accurate fingerprint predictions typically worsens retrieval results, and vice versa. Our theoretical analysis shows this trade-off depends on the similarity structure of candidate sets, providing guidance for loss function and fingerprint selection.

质谱分析深度学习指纹预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。