用扩散模型提升肽段从头测序准确率,突破传统序列生成局限。
Diffusion Decoding for Peptide De Novo Sequencing
- 引入扩散解码器替代自回归模型,可从任意片段开始生成序列。
- 最佳设计下氨基酸召回率较基线提升0.373,统计显著。
- 适合蛋白质组学、质谱数据分析方向研究者参考。
肽段从头测序用于在不依赖现有蛋白数据库的情况下,从串联质谱数据中重构氨基酸序列。传统深度学习方法如Casanovo主要采用自回归解码器,逐个预测氨基酸,易产生累积错误且难以利用高置信度区域。本文探索适配离散数据域的扩散解码器,其允许从任意肽段片段开始生成,提升预测精度。实验对比三种扩散解码器设计、背包束搜索及不同损失函数。结果表明,背包束搜索未提升性能,单纯替换变压器解码器反而降低表现。尽管肽段精确率和召回率仍为0,但最优扩散解码器搭配DINOISER损失函数时,氨基酸召回率相较基线自回归模型提升0.373,具有统计显著性。该结果表明扩散解码器不仅能提高模型敏感性,还可推动肽段从头测序技术取得实质性进展。
原文摘要 · Abstract (English)
Peptide de novo sequencing is a method used to reconstruct amino acid sequences from tandem mass spectrometry data without relying on existing protein sequence databases. Traditional deep learning approaches, such as Casanovo, mainly utilize autoregressive decoders and predict amino acids sequentially. Subsequently, they encounter cascading errors and fail to leverage high-confidence regions effectively. To address these issues, this paper investigates using diffusion decoders adapted for the discrete data domain. These decoders provide a different approach, allowing sequence generation to start from any peptide segment, thereby enhancing prediction accuracy. We experiment with three different diffusion decoder designs, knapsack beam search, and various loss functions. We find knapsack beam search did not improve performance metrics and simply replacing the transformer decoder with a diffusion decoder lowered performance. Although peptide precision and recall were still 0, the best diffusion decoder design with the DINOISER loss function obtained a statistically significant improvement in amino acid recall by 0.373 compared to the baseline autoregressive decoder-based Casanovo model. These findings highlight the potential of diffusion decoders to not only enhance model sensitivity but also drive significant advancements in peptide de novo sequencing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。