用化学语言模型提升质谱分子识别的跨仪器泛化能力
Contrastive Domain Generalization for Cross-Instrument Molecular Identification in Mass Spectrometry
- 将质谱图直接映射到预训练化学语言模型的结构嵌入空间
- 在256类零样本检索中达到42.2%的准确率,5类5次少样本重识别达95.4%
- 适合需要跨仪器、跨数据集泛化的分子鉴定研究者
从质谱(MS)数据中识别分子仍面临物理谱峰与化学结构之间语义鸿沟的根本挑战。现有深度学习方法常将光谱匹配视为闭集识别任务,限制了对未见分子骨架的泛化能力。为此,我们提出一种跨模态对齐框架,直接将质谱图映射至预训练化学语言模型的化学结构嵌入空间。在严格的骨架不交集基准测试中,该模型在固定256类零样本检索中达到42.2%的Top-1准确率,并在全局检索设置下展现出强泛化能力。此外,学习到的嵌入空间具有强化学一致性,在5类5次少样本分子重识别任务中达到95.4%准确率。结果表明,显式融合物理谱图分辨率与分子结构嵌入是解决质谱分子识别泛化瓶颈的关键。
原文摘要 · Abstract (English)
Identifying molecules from mass spectrometry (MS) data remains a fundamental challenge due to the semantic gap between physical spectral peaks and underlying chemical structures. Existing deep learning approaches often treat spectral matching as a closed-set recognition task, limiting their ability to generalize to unseen molecular scaffolds. To overcome this limitation, we propose a cross-modal alignment framework that directly maps mass spectra into the chemically meaningful molecular structure embedding space of a pretrained chemical language model. On a strict scaffold-disjoint benchmark, our model achieves a Top-1 accuracy of 42.2% in fixed 256-way zero-shot retrieval and demonstrates strong generalization under a global retrieval setting. Moreover, the learned embedding space demonstrates strong chemical coherence, reaching 95.4% accuracy in 5-way 5-shot molecular re-identification. These results suggest that explicitly integrating physical spectral resolution with molecular structure embedding is key to solving the generalization bottleneck in molecular identification from MS data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。