arXiv:2510.26715cs.LG2025-10被引 1

用大规模模型提升质谱数据识别与生物解读能力

LSM-MS2: A Foundation Model Bridging Spectral Identification and Biological Interpretation

  • 基于百万级质谱数据训练深度学习模型,构建化学语义空间
  • 对同分异构体识别准确率提升30%,复杂样本中正确识别率提高42%
  • 可直接从少量数据中实现疾病状态区分和临床预测

绝大多数质谱数据尚未被充分解析,其中蕴含的生物与化学信息未被挖掘。近年来机器学习进展开始弥补这一缺口,尤其在串联质谱数据的谱图识别任务中表现突出。本文提出 LSM-MS2 的最新版本,一种在数百万条谱图上训练的大规模深度学习基础模型,旨在学习一个语义化的化学空间。该模型在谱图识别任务中达到当前最佳性能,对具有挑战性的同分异构体识别准确率提升30%,在复杂生物样本中正确识别数量增加42%,且在低浓度条件下仍保持鲁棒性。此外,LSM-MS2 生成丰富的谱图嵌入,可直接用于生物解读,仅需少量下游数据即可成功区分疾病状态,并在多种转化医学应用中预测临床结果。

原文摘要 · Abstract (English)

A vast majority of mass spectrometry data remains uncharacterized, leaving much of its biological and chemical information untapped. Recent advances in machine learning have begun to address this gap, particularly for tasks such as spectral identification in tandem mass spectrometry data. Here, we present the latest generation of LSM-MS2, a large-scale deep learning foundation model trained on millions of spectra to learn a semantic chemical space. LSM-MS2 achieves state-of-the-art performance in spectral identification, improving on existing methods by 30% in accuracy of identifying challenging isomeric compounds, yielding 42% more correct identifications in complex biological samples, and maintaining robustness under low-concentration conditions. Furthermore, LSM-MS2 produces rich spectral embeddings that enable direct biological interpretation from minimal downstream data, successfully differentiating disease states and predicting clinical outcomes across diverse translational applications.

质谱分析深度学习生物标志物基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。