用多模型联合重排序提升肽段从头测序精度。
Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing
- 基于多个模型候选结果,用列表级重排序优化肽段序列。
- 在多个数据集上超越基线模型,零样本泛化能力出色。
- 适合需要高精度肽段鉴定的蛋白质组学研究者。
从头肽段测序是蛋白质组学中的关键任务。现有深度学习方法受限于质谱数据的复杂性及噪声分布异质性,易产生数据特定偏差。本文提出RankNovo,首个利用多模型互补优势的深度重排序框架。RankNovo采用列表级重排序策略,将候选肽段建模为多重序列比对,并通过轴向注意力提取跨候选序列的特征。同时引入两个新指标:PMD(肽段质量偏差)和RMD(残基剩余质量偏差),在序列与残基层面量化质量差异,提供精细监督信号。大量实验表明,RankNovo不仅超越用于生成训练候选的基线模型,还创下新的性能基准。更重要的是,其在未见过的模型生成结果上表现出强零样本泛化能力,验证了框架的鲁棒性与通用性。本工作提出一种颠覆单模型范式的新型重排序策略,推动从头肽段测序精度边界。代码已开源。
原文摘要 · Abstract (English)
De novo peptide sequencing is a critical task in proteomics. However, the performance of current deep learning-based methods is limited by the inherent complexity of mass spectrometry data and the heterogeneous distribution of noise signals, leading to data-specific biases. We present RankNovo, the first deep reranking framework that enhances de novo peptide sequencing by leveraging the complementary strengths of multiple sequencing models. RankNovo employs a list-wise reranking approach, modeling candidate peptides as multiple sequence alignments and utilizing axial attention to extract informative features across candidates. Additionally, we introduce two new metrics, PMD (Peptide Mass Deviation) and RMD (residual Mass Deviation), which offer delicate supervision by quantifying mass differences between peptides at both the sequence and residue levels. Extensive experiments demonstrate that RankNovo not only surpasses its base models used to generate training candidates for reranking pre-training, but also sets a new state-of-the-art benchmark. Moreover, RankNovo exhibits strong zero-shot generalization to unseen models whose generations were not exposed during training, highlighting its robustness and potential as a universal reranking framework for peptide sequencing. Our work presents a novel reranking strategy that fundamentally challenges existing single-model paradigms and advances the frontier of accurate de novo sequencing. Our source code is provided on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。