arXiv:2510.23746cs.AIcs.LG2025-10被引 4

用测试时调优让语言模型直接从质谱图生成分子结构,无需中间步骤。

Test-Time Tuned Language Models Enable End-to-end De Novo Molecular Structure Generation from MS/MS Spectra

  • 基于Transformer的端到端框架,输入质谱图和分子式直接生成结构。
  • 在MassSpecGym上达到3.16%的Top-1准确率,NPLIB1上达12.88%。
  • 测试时调优提升泛化能力,适合处理未见谱图的科研场景。

串联质谱是代谢组学、天然产物发现和环境分析中鉴定未知小分子的核心技术。然而,其碎片化过程的不确定性及化学空间庞大,使结构解析极具挑战性,尤其当部署与训练条件不一致时。现有方法依赖数据库匹配已知分子谱图或需多步流程,包括中间指纹预测或昂贵的片段标注。本文提出一种基于Transformer的端到端框架,直接根据输入的串联质谱图及其分子式生成分子结构,避免人工标注和中间步骤,同时利用模拟数据进行迁移学习。为应对分布外谱图问题,引入测试时调优策略,动态适应新实验数据。该方法在MassSpecGym基准上取得3.16%的Top-1准确率,在NPLIB1数据集上达12.88%,显著优于传统微调。基线方法分别被超越27%和67%。即使未恢复确切参考结构,生成候选仍具高化学合理性,体现强刘易斯相似度(平均提升83%于NPLIB1,64%于MassSpecGym)。该框架兼具简洁性与适应性,可为专家解析未知谱图提供有效指导。

原文摘要 · Abstract (English)

Tandem Mass Spectrometry is a cornerstone technique for identifying unknown small molecules in fields such as metabolomics, natural product discovery and environmental analysis. However, certain aspects, such as the probabilistic fragmentation process and size of the chemical space, make structure elucidation from such spectra highly challenging, particularly when there is a shift between the deployment and training conditions. Current methods rely on database matching of previously observed spectra of known molecules and multi-step pipelines that require intermediate fingerprint prediction or expensive fragment annotations. We introduce a novel end-to-end framework based on a transformer model that directly generates molecular structures from an input tandem mass spectrum and its corresponding molecular formula, thereby eliminating the need for manual annotations and intermediate steps, while leveraging transfer learning from simulated data. To further address the challenge of out-of-distribution spectra, we introduce a test-time tuning strategy that dynamically adapts the pre-trained model to novel experimental data. Our approach achieves a Top-1 accuracy of 3.16% on the MassSpecGym benchmark and 12.88% on the NPLIB1 datasets, considerably outperforming conventional fine-tuning. Baseline approaches are also surpassed by 27% and 67% respectively. Even when the exact reference structure is not recovered, the generated candidates are chemically informative, exhibiting high structural plausibility as reflected by strong Tanimoto similarity to the ground truth. Notably, we observe a relative improvement in average Tanimoto similarity of 83% on NPLIB1 and 64% on MassSpecGym compared to state-of-the-art methods. Our framework combines simplicity with adaptability, generating accurate molecular candidates that offer valuable guidance for expert interpretation of unseen spectra.

分子生成质谱解析测试时调优端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。