提出新方法让质谱肽段测序更忠实于原始数据,提升准确性。
MemNovo: Look Back at the Spectrum for Balanced De Novo Peptide Sequencing from Mass Spectrometry

- 通过构建持续的谱图记忆库,在推理时重新平衡肽段与谱图信息
- 在九物种基准上使肽段精度最高提升39.1%,计算开销极低
- 无需训练即可插入现有模型,适合做质谱分析的研究者
从串联质谱进行从头肽段测序是蛋白质组学的关键,可识别无参考数据库的新肽段。尽管近期基于Transformer的编码器-解码器模型表现优异,但我们发现其推理过程存在关键缺陷:自回归解码器过度依赖生成序列先验,逐步忽略输入质谱中的细粒度物理证据。这导致生成的肽段虽生物合理,却与原始谱图不一致。为此,我们提出MemNovo,一种无需训练、即插即用的推理阶段调衡机制。该方法通过建立持久的谱图记忆库,将检索到的特征以极保守的残差连接注入最终解码阶段。理论分析表明,此机制恢复了解码状态与原始谱图间的互信息。在九物种基准上对两种代表性基线模型Casanovo和InstaNovo的大量实验显示,MemNovo持续提升氨基酸精度与肽段精度,对Casanovo最高实现39.1%的肽段精度相对提升,对InstaNovo提升3.9%,且计算开销可忽略。
原文摘要 · Abstract (English)
De novo peptide sequencing from tandem mass spectrometry is pivotal in proteomics, enabling identification of novel peptides without reference databases. While recent Transformer-based encoder-decoder models have achieved remarkable performance, we uncover a critical pathology in their inference dynamics. Through comprehensive feature scaling experiments, we demonstrate that existing auto-regressive peptide decoders tend to over-rely on generated-sequence priors while progressively under-utilizing fine-grained physical evidence from the input mass spectrum. This phenomenon leads to suboptimal results, where generated peptide sequences are biologically plausible yet not faithful to the input spectrum. To rectify this, we propose MemNovo, a training-free and plug-and-play mechanism that re-balances peptide and spectral contributions at inference time. MemNovo alleviates the information bottleneck by establishing a persistent spectral memory bank and injecting retrieved features directly into the final decoding stage via an ultra-conservative residual connection. Theoretical analysis confirms that this mechanism restores the mutual information between the decoder state and the raw spectrum. Extensive experiments on the Nine Species benchmark with two representative baselines, Casanovo and InstaNovo, demonstrate that MemNovo consistently improves both amino acid precision and peptide precision, achieving up to 39.1% relative improvement in peptide precision for Casanovo and up to 3.9% for InstaNovo, with negligible computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。