用生成模型提升质谱图查分子结构的准确率
Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra
- 分两阶段:先用对比学习找候选分子,再用生成模型优化结构
- 在两个数据集上顶1准确率提升超40%,泛化能力更强
- 适合需要精准化合物识别的科研与药物研发人员
从串联质谱图中检索分子结构是快速化合物鉴定的关键步骤。现有方法如传统质谱库匹配存在库覆盖有限的问题,而近期跨模态表示学习框架常因模态错位导致检索精度与泛化能力不足。为此,我们提出GLMR——一种基于生成语言模型的检索框架,通过两阶段流程缓解跨模态错位问题。预检索阶段采用基于对比学习的模型,为输入质谱图识别出前若干候选分子作为上下文先验;生成检索阶段将这些候选分子与输入质谱结合,引导生成模型生成更优分子结构,并基于分子相似性对候选进行重排序。在MassSpecGym和新提出的MassRET-20k数据集上的实验表明,GLMR显著优于现有方法,顶1准确率提升超过40%,且表现出强泛化能力。
原文摘要 · Abstract (English)
Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。