arXiv:2601.17204cs.LGcs.CE2026-01被引 2

用几何对齐法提升质谱分子识别准确率

SpecBridge: Bridging Mass Spectrometry and Molecular Representations via Cross-Modal Alignment

  • 将质谱与分子表征通过隐式对齐映射到同一空间
  • 在三大基准上相对强基线提升20%-25%准确率
  • 仅微调少量参数,适合资源受限场景

非靶向小分子质谱(MS/MS)鉴定仍面临谱图库不全的瓶颈。现有深度学习方法多为两种极端:显式生成模型逐原子构建分子图,或从零学习跨模态子空间的对比模型。本文提出SpecBridge,一种新型隐式对齐框架,将结构识别视为几何对齐问题。该方法微调自监督谱图编码器DreaMS,直接投影至冻结的分子基础模型ChemBERTa的潜在空间,并通过余弦相似度检索预计算的分子嵌入库。在MassSpecGym、Spectraverse和MSnLib三大基准上,SpecBridge相较强基线实现约20%-25%的顶1检索准确率提升,同时保持可训练参数极少。结果表明,对齐冻结基础模型是一种稳定且实用的替代方案。代码已开源:https://github.com/HassounLab/SpecBridge。

原文摘要 · Abstract (English)

Small-molecule identification from tandem mass spectrometry (MS/MS) remains a bottleneck in untargeted settings where spectral libraries are incomplete. While deep learning offers a solution, current approaches typically fall into two extremes: explicit generative models that construct molecular graphs atom-by-atom, or joint contrastive models that learn cross-modal subspaces from scratch. We introduce SpecBridge, a novel implicit alignment framework that treats structure identification as a geometric alignment problem. SpecBridge fine-tunes a self-supervised spectral encoder (DreaMS) to project directly into the latent space of a frozen molecular foundation model (ChemBERTa), and then performs retrieval by cosine similarity to a fixed bank of precomputed molecular embeddings. Across MassSpecGym, Spectraverse, and MSnLib benchmarks, SpecBridge improves top-1 retrieval accuracy by roughly 20-25% relative to strong neural baselines, while keeping the number of trainable parameters small. These results suggest that aligning to frozen foundation models is a practical, stable alternative to designing new architectures from scratch. The code for SpecBridge is released at https://github.com/HassounLab/SpecBridge.

质谱分析跨模态对齐分子表示检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。