arXiv:2602.01772cs.LGcs.AI2026-02

DIA-CLIP让质谱分析无需训练就能精准识别蛋白质,跨物种通用。

DIA-CLIP: a universal representation learning framework for zero-shot DIA proteomics

  • 用对比学习+编码器解码器构建肽段与谱图的统一表征
  • 零样本下蛋白识别率提升45%,共检出物减少12%
  • 适合单细胞、空间蛋白组学等高难度应用

数据非依赖采集质谱(DIA-MS)已成为蛋白质组学和系统生物学的重要工具,具有深度和可重复性优势。然而现有分析框架需每批次进行半监督训练以重评分肽段-谱图匹配(PSM),易过拟合并缺乏跨物种和实验条件的泛化能力。本文提出DIA-CLIP,一种预训练模型,将DIA分析范式从半监督训练转向通用的跨模态表征学习。通过融合双编码器对比学习与编码器-解码器架构,DIA-CLIP建立肽段与对应谱图特征的统一表示,实现高精度零样本PSM推断。在多个基准测试中,DIA-CLIP持续优于现有先进工具,蛋白识别率最高提升45%,共检出物识别降低12%。该模型在单细胞和空间蛋白组学等实际场景中潜力巨大,可助力新生物标志物发现及复杂细胞机制解析。

原文摘要 · Abstract (English)

Data-independent acquisition mass spectrometry (DIA-MS) has established itself as a cornerstone of proteomic profiling and large-scale systems biology, offering unparalleled depth and reproducibility. Current DIA analysis frameworks, however, require semi-supervised training within each run for peptide-spectrum match (PSM) re-scoring. This approach is prone to overfitting and lacks generalizability across diverse species and experimental conditions. Here, we present DIA-CLIP, a pre-trained model shifting the DIA analysis paradigm from semi-supervised training to universal cross-modal representation learning. By integrating dual-encoder contrastive learning framework with encoder-decoder architecture, DIA-CLIP establishes a unified cross-modal representation for peptides and corresponding spectral features, achieving high-precision, zero-shot PSM inference. Extensive evaluations across diverse benchmarks demonstrate that DIA-CLIP consistently outperforms state-of-the-art tools, yielding up to a 45% increase in protein identification while achieving a 12% reduction in entrapment identifications. Moreover, DIA-CLIP holds immense potential for diverse practical applications, such as single-cell and spatial proteomics, where its enhanced identification depth facilitates the discovery of novel biomarkers and the elucidates of intricate cellular mechanisms.

质谱分析零样本学习蛋白质组学跨模态表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。