arXiv:2505.10848cs.LG2025-05被引 4

用自编码器统一建模质谱数据,提升蛋白质组学分析性能。

Foundation model for mass spectrometry proteomics

  • 以从头测序为预训练任务,构建通用质谱表示模型。
  • 在质量、嵌合、磷酸化等4项下游任务上均提升性能。
  • 适合缺乏标注数据的质谱分析场景,尤其利于新任务迁移。

质谱技术是蛋白质组学领域主流方法,可高通量分析复杂生物样本中的蛋白质组成。由于仪器复杂且数据量大,需借助复杂的计算方法处理与解读质谱数据。机器学习在提升质谱数据分析方面展现出巨大潜力,已有众多针对特定流程环节优化的方法被广泛应用。本文提出将多种质谱预测任务统一于一个质谱基础模型下。我们以从头测序作为预训练任务,对质谱编码器进行预训练,并验证了该预训练表示在四个下游任务(质谱质量预测、嵌合性预测、磷酸化预测、糖基化状态预测)上的性能提升。进一步通过多任务微调,发现各任务独立性能均得到改善。结果表明,基于从头测序训练的串联质谱蛋白质组学基础模型能学习到可泛化的质谱表征,在训练数据有限的下游任务中表现更优,有望提升蛋白质组学实验的数据采集与分析效率。

原文摘要 · Abstract (English)

Mass spectrometry is the dominant technology in the field of proteomics, enabling high-throughput analysis of the protein content of complex biological samples. Due to the complexity of the instrumentation and resulting data, sophisticated computational methods are required for the processing and interpretation of acquired mass spectra. Machine learning has shown great promise to improve the analysis of mass spectrometry data, with numerous purpose-built methods for improving specific steps in the data acquisition and analysis pipeline reaching widespread adoption. Here, we propose unifying various spectrum prediction tasks under a single foundation model for mass spectra. To this end, we pre-train a spectrum encoder using de novo sequencing as a pre-training task. We then show that using these pre-trained spectrum representations improves our performance on the four downstream tasks of spectrum quality prediction, chimericity prediction, phosphorylation prediction, and glycosylation status prediction. Finally, we perform multi-task fine-tuning and find that this approach improves the performance on each task individually. Overall, our work demonstrates that a foundation model for tandem mass spectrometry proteomics trained on de novo sequencing learns generalizable representations of spectra, improves performance on downstream tasks where training data is limited, and can ultimately enhance data acquisition and analysis in proteomics experiments.

质谱分析基础模型蛋白质组学多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。