用深度学习直接从质谱数据解码蛋白质修饰,发现更多未知蛋白变化。
Accurate de novo sequencing of the modified proteome with OmniNovo
- 基于深度学习的统一模型,无需参考数据库直接解析肽段序列。
- 在1%错误发现率下比传统方法多识别51%的肽段。
- 能发现训练时未见的修饰位点,适合探索未知蛋白质调控机制。
翻译后修饰(PTMs)是调控蛋白质功能的动态化学语言,但现有蛋白质组学方法仍无法检测大量修饰蛋白。标准数据库搜索算法因搜索空间呈组合爆炸式增长,难以识别未表征或复杂修饰。本文提出OmniNovo,一种统一的深度学习框架,可直接从串联质谱中进行无参考的未修饰与修饰肽段从头测序。与仅限特定修饰类型工具不同,OmniNovo学习通用断裂规律,在单一模型中解析多种修饰。通过结合质量约束解码算法与严格的错误发现率估计,OmniNovo达到当前最佳准确率,在1%错误发现率下比标准方法多识别51%的肽段。关键的是,该模型可泛化至训练时未见的生物位点,揭示了蛋白质组中的“暗物质”,实现细胞调控的无偏全面分析。
原文摘要 · Abstract (English)
Post-translational modifications (PTMs) serve as a dynamic chemical language regulating protein function, yet current proteomic methods remain blind to a vast portion of the modified proteome. Standard database search algorithms suffer from a combinatorial explosion of search spaces, limiting the identification of uncharacterized or complex modifications. Here we introduce OmniNovo, a unified deep learning framework for reference-free sequencing of unmodified and modified peptides directly from tandem mass spectra. Unlike existing tools restricted to specific modification types, OmniNovo learns universal fragmentation rules to decipher diverse PTMs within a single coherent model. By integrating a mass-constrained decoding algorithm with rigorous false discovery rate estimation, OmniNovo achieves state-of-the-art accuracy, identifying 51\% more peptides than standard approaches at a 1\% false discovery rate. Crucially, the model generalizes to biological sites unseen during training, illuminating the dark matter of the proteome and enabling unbiased comprehensive analysis of cellular regulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。