为质谱分子结构检索设计可信度评估机制,只信任高置信预测。
When should we trust the annotation? Selective prediction for molecular structure retrieval from mass spectra
- 基于多层级不确定性量化,区分可信与不可信预测
- 检索级总不确定性是最佳拒绝标准,可控制错误率
- 无需假设分布,直接保证误差率在可控范围内
从串联质谱(MS/MS)中识别分子结构的机器学习方法发展迅速,但依然存在显著错误率。在临床代谢组学和环境筛查等高风险场景中,错误标注可能导致严重后果,因此必须判断何时可信赖预测结果。本文提出一种选择性预测框架,将低风险预测自动分离出低置信度预测。在风险-覆盖权衡框架下,系统评估了三类不确定性量化策略:学习表示空间中的输入级距离、分子指纹位的指纹级不确定性、以及候选排序的检索级不确定性。比较了包括一阶置信度、二阶分布下的随机与认知不确定性估计、以及潜在空间距离在内的多种评分函数。所有实验均在MassSpecGym基准上进行。分析表明,指纹级不确定性不适合作为检索成功性的代理指标,而检索级总不确定性提供最强的拒绝准则;一阶置信度虽简单,但仍是有效的基线。我们证明,通过基于泛化界的风险控制,从业者可指定可接受的错误率,并获得以高概率满足该约束的预测子集。
原文摘要 · Abstract (English)
Machine learning methods for identifying molecular structures from tandem mass spectra (MS/MS) have advanced rapidly, yet current approaches still exhibit significant error rates. In high-stakes applications such as clinical metabolomics and environmental screening, incorrect annotations can have serious consequences, making it essential to determine when a prediction can be trusted. In this work, we introduce a selective prediction framework for molecular structure retrieval from MS/MS spectra, separating low-risk predictions automatically from lower-confidence predictions. We formulate the problem within the risk-coverage tradeoff framework and systematically evaluate uncertainty quantification strategies at three levels: input-level distance in the learned representation, fingerprint-level uncertainty over predicted molecular fingerprint bits, and retrieval-level uncertainty over candidate rankings. We compare scoring functions including first-order confidence measures, aleatoric and epistemic uncertainty estimates from second-order distributions, as well as distance-based measures in the latent space. All experiments are conducted on the MassSpecGym benchmark. Our analysis reveals that while fingerprint-level uncertainty scores are poor proxies for retrieval success, retrieval-level total uncertainty provides the overall strongest rejection criterion, and first-order confidence measures are computationally inexpensive, strong baselines. We demonstrate that by applying distribution-free risk control via generalisation bounds, practitioners can specify a tolerable error rate and obtain a subset of annotations satisfying that constraint with high probability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。