arXiv:2605.01945cs.LGcs.AI2026-05

构建统一基准,提升肽段质谱预测的公平性与可靠性

PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction

论文配图:PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
图 1 · 摘自论文原文
  • 整合多源数据并统一预处理,避免模型比较偏差
  • 采用骨架不重叠划分策略,杜绝序列泄露问题
  • 跨物种评估+实验扰动测试,揭示模型真实鲁棒性

串联质谱为复杂生物样本中蛋白质的高通量鉴定与定量提供了框架。在计算蛋白组学中,预测肽段的质谱/质谱(MS/MS)谱图是关键任务,支撑大规模肽段鉴定与定量。尽管深度学习显著提升了预测精度,但三个评估挑战掩盖了领域真实进展:一是数据预处理不一致、输出空间不兼容,阻碍公平比较;二是数据划分策略缺陷导致隐式序列泄露,虚高性能报告;三是现有评估缺乏跨物种系统性测试和对关键实验条件影响的鲁棒性分析。为此,我们提出PepSpecBench,一个统一的肽段MS/MS谱图预测基准。该基准统一多个互补公开数据集的预处理流程,采用严格的骨架不重叠划分策略消除序列泄露,并在共享的碎片离子表示空间内评估多种架构。进一步引入全面的多物种评估套件及基于物理的元数据扰动探针,以检验模型鲁棒性与仪器感知能力。我们发现了六种代表性模型间先前未被识别的性能差异与鲁棒性局限,为未来模型设计、评估与实际部署提供可操作洞见。

原文摘要 · Abstract (English)

Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational proteomics, predicting peptide MS/MS spectra is a critical task, enabling downstream applications such as large-scale peptide identification and quantification. While deep learning architectures have substantially improved prediction accuracy, three evaluation challenges obscure the true progress of the field. First, inconsistent data preprocessing and incompatible model output spaces hinder fair model comparison. Second, flawed data splitting strategies can permit hidden sequence leakage and inflate reported performance. Third, existing evaluations typically lack comprehensive cross-species benchmarking and systematic assessment of model robustness to influential experimental conditions. To address these challenges, we propose PepSpecBench, a unified benchmark for peptide MS/MS spectrum prediction. PepSpecBench standardizes data preprocessing across complementary public datasets, enforces a strict backbone-disjoint splitting strategy to eliminate sequence leakage, and evaluates diverse architectures within a shared fragment-ion representation space. It further introduces a comprehensive multi-species evaluation suite and physically grounded metadata perturbation probes to assess model robustness and instrument awareness. We uncover previously unrecognized performance discrepancies and robustness limitations across six representative models, providing actionable insights for future model design, evaluation and practical deployment.

质谱预测基准测试蛋白组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。