arXiv:2510.25814q-bio.QMcs.LG2025-10中稿 · BIBM 2025被引 1

通过预测镜像肽键断裂优化序列设计,提升生物数据存储的可读性。

Optimizing Mirror-Image Peptide Sequence Design for Data Storage via Peptide Bond Cleavage Prediction

  • 用深度学习模型DBond预测镜像肽键断裂位置,指导易测序序列设计。
  • 构建513个镜像肽的质谱数据集,生成约1250万条标签数据。
  • 提出双策略预测方法,显著提升序列解析准确率,适合生物存储研究者。

传统非生物存储介质如硬盘在大数据时代面临存储密度和寿命的双重瓶颈。由D-氨基酸组成的镜像肽因其高密度、结构稳定和长寿命,成为有前景的生物存储介质。镜像肽序列依赖从头合成技术,但其准确性受限于串联质谱数据稀缺以及现有算法处理此类肽的困难。本研究首次提出通过优化镜像肽序列设计来间接提升测序精度。我们构建了包含513个镜像肽的MiPD513质谱数据集,开发了肽键断裂标注算法(PBCLA),基于该数据集生成约1250万条标注数据。提出结合多标签与单标签分类的双重预测策略。在独立测试集上,单标签分类策略在单键与多键断裂预测任务中均优于其他方法,为序列优化提供了坚实基础。

原文摘要 · Abstract (English)

Traditional non-biological storage media, such as hard drives, face limitations in both storage density and lifespan due to the rapid growth of data in the big data era. Mirror-image peptides composed of D-amino acids have emerged as a promising biological storage medium due to their high storage density, structural stability, and long lifespan. The sequencing of mirror-image peptides relies on \textit{de-novo} technology. However, its accuracy is limited by the scarcity of tandem mass spectrometry datasets and the challenges that current algorithms encounter when processing these peptides directly. This study is the first to propose improving sequencing accuracy indirectly by optimizing the design of mirror-image peptide sequences. In this work, we introduce DBond, a deep neural network based model that integrates sequence features, precursor ion properties, and mass spectrometry environmental factors for the prediction of mirror-image peptide bond cleavage. In this process, sequences with a high peptide bond cleavage ratio, which are easy to sequence, are selected. The main contributions of this study are as follows. First, we constructed MiPD513, a tandem mass spectrometry dataset containing 513 mirror-image peptides. Second, we developed the peptide bond cleavage labeling algorithm (PBCLA), which generated approximately 12.5 million labeled data based on MiPD513. Third, we proposed a dual prediction strategy that combines multi-label and single-label classification. On an independent test set, the single-label classification strategy outperformed other methods in both single and multiple peptide bond cleavage prediction tasks, offering a strong foundation for sequence optimization.

生物存储镜像肽深度学习质谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。