构建首个肽特异性碎片离子概率预测基准,提升质谱蛋白组学识别准确率。
Pep2Prob Benchmark: Predicting Fragment Ion Probability for MS$^2$-based Proteomics
- 基于60万+肽段的高分辨率质谱数据,建立肽特异性碎片概率预测基准
- 引入肽段序列与电荷态信息,显著优于传统全局统计方法
- 适合从事质谱分析、蛋白质组学与机器学习交叉研究者参考
蛋白质在几乎所有的细胞功能中发挥作用,并构成绝大多数药物靶点,其分析对理解健康与疾病中的人类生物学至关重要。串联质谱(MS²)是蛋白质组学的主要分析技术,通过电离肽段、碎裂并解析所得质谱图来鉴定和定量生物样本中的蛋白质。在MS²分析中,肽段碎片离子概率预测起关键作用,可补充强度信息以提高肽段识别精度。现有方法依赖碎片化的全局统计,假设某碎片概率在所有肽段中均一,但这一假设过于简化,且不符合生化原理,限制了预测准确性。为此,我们提出Pep2Prob,首个专为肽段特异性碎片离子概率预测设计的综合数据集与基准。该数据集包含608,780个唯一前体(每个前体对应一条肽段序列与电荷态),源自超过18300万条高质量、高分辨率的HCD MS²谱图,且具有验证的肽段归属与碎裂标注。我们采用简单统计规则与基于学习的方法建立基线性能,发现利用肽段特异性信息的模型显著优于仅依赖全局碎片统计的方法。此外,随着模型容量增加,性能持续提升,表明肽段-碎裂关系存在复杂非线性特征,需采用先进的机器学习方法建模。
原文摘要 · Abstract (English)
Proteins perform nearly all cellular functions and constitute most drug targets, making their analysis fundamental to understanding human biology in health and disease. Tandem mass spectrometry (MS$^2$) is the major analytical technique in proteomics that identifies peptides by ionizing them, fragmenting them, and using the resulting mass spectra to identify and quantify proteins in biological samples. In MS$^2$ analysis, peptide fragment ion probability prediction plays a critical role, enhancing the accuracy of peptide identification from mass spectra as a complement to the intensity information. Current approaches rely on global statistics of fragmentation, which assumes that a fragment's probability is uniform across all peptides. Nevertheless, this assumption is oversimplified from a biochemical principle point of view and limits accurate prediction. To address this gap, we present Pep2Prob, the first comprehensive dataset and benchmark designed for peptide-specific fragment ion probability prediction. The proposed dataset contains fragment ion probability statistics for 608,780 unique precursors (each precursor is a pair of peptide sequence and charge state), summarized from more than 183 million high-quality, high-resolution, HCD MS$^2$ spectra with validated peptide assignments and fragmentation annotations. We establish baseline performance using simple statistical rules and learning-based methods, and find that models leveraging peptide-specific information significantly outperform previous methods using only global fragmentation statistics. Furthermore, performance across benchmark models with increasing capacities suggests that the peptide-fragmentation relationship exhibits complex nonlinearities requiring sophisticated machine learning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。