首个统一谱图解析的深度学习模型,显著提升肽段识别率与修饰覆盖。
pUniFind: a unified large pre-trained deep learning model pushing the limit of mass spectra interpretation
- 通过跨模态预训练统一肽段与谱图匹配,实现端到端评分与零样本从头测序。
- 在免疫肽组学中肽段鉴定数提升42.6%,修饰种类超1300种,鉴定率高出60%。
- 支持大规模搜索空间,具备自检能力,适合高通量蛋白质组学研究者使用。
深度学习推动质谱数据分析进步,但多数模型仍为特征提取器而非统一评分框架。我们提出pUniFind,首个在蛋白质组学中实现大规模多模态预训练的统一模型,集成端到端肽段-谱图打分与开放零样本从头测序。模型在超过一亿条开放搜索获得的谱图上训练,通过跨模态预测对齐谱图与肽段模态,在多种数据集上优于传统引擎,尤其在免疫肽组学中肽段鉴定数提升42.6%。支持超过1,300种修饰,尽管搜索空间扩大300倍,仍比现有从头测序方法多鉴定60%的PSM。基于深度学习的质量控制模块进一步恢复38.5%额外肽段,包括1,891个映射至基因组但未在参考蛋白组中出现的肽段,同时保持完整碎片离子覆盖。这些结果确立了一个统一、可扩展的深度学习分析框架,显著提升敏感性、修饰覆盖与可解释性。
原文摘要 · Abstract (English)
Deep learning has advanced mass spectrometry data interpretation, yet most models remain feature extractors rather than unified scoring frameworks. We present pUniFind, the first large-scale multimodal pre-trained model in proteomics that integrates end-to-end peptide-spectrum scoring with open, zero-shot de novo sequencing. Trained on over 100 million open search-derived spectra, pUniFind aligns spectral and peptide modalities via cross modality prediction and outperforms traditional engines across diverse datasets, particularly achieving a 42.6 percent increase in the number of identified peptides in immunopeptidomics. Supporting over 1,300 modifications, pUniFind identifies 60 percent more PSMs than existing de novo methods despite a 300-fold larger search space. A deep learning based quality control module further recovers 38.5 percent additional peptides including 1,891 mapped to the genome but absent from reference proteomes while preserving full fragment ion coverage. These results establish a unified, scalable deep learning framework for proteomic analysis, offering improved sensitivity, modification coverage, and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。