用结构化提示让大模型精准模拟质谱碎片,提升代谢物注释效率。
General Intelligence-based Fragmentation (GIF): A framework for peak-labeled spectra simulation
- 通过标签、结构化输入输出和迭代优化,引导大模型进行质谱模拟
- GPT-4o在真实与模拟谱图间达到0.36的余弦相似度,优于其他模型
- 适合需要可解释推理的代谢组学研究者,支持人机协作流程
尽管参考库日益丰富且计算工具进步,代谢组学中谱图注释率仍较低。大型语言模型(LLMs)在生成与推理任务中表现优异,激发其在质谱注释等科学问题中的应用。本文提出通用智能驱动的碎片化框架(General Intelligence-based Fragmentation, GIF),通过结构化提示与推理引导预训练模型进行谱图模拟。GIF结合标签、结构化输入/输出、系统提示、指令式提示及迭代优化,提供系统性替代传统随意提示的方法。我们基于MassSpecGym数据集构建了新的问答数据集MassSpecGym QA-sim,评估当前通用大模型在碎片化推理与强度预测方面的性能。微调后,GPT-4o与GPT-4o-mini在模拟谱图与真实谱图间分别取得0.36和0.35的余弦相似度,优于GPT-5、Llama-3.1与ChemDFM等模型。GIF还超越多个深度学习基线。结果表明,该框架不仅适用于谱图模拟,还能推动人机协同工作与可解释的分子碎片化推理。
原文摘要 · Abstract (English)
Despite growing reference libraries and advanced computational tools, progress in the field of metabolomics remains constrained by low rates of annotating measured spectra. The recent developments of large language models (LLMs) have led to strong performance across a wide range of generation and reasoning tasks, spurring increased interest in LLMs' application to domain-specific scientific challenges, such as mass spectra annotation. Here, we present a novel framework, General Intelligence-based Fragmentation (GIF), that guides pretrained LLMs through spectra simulation using structured prompting and reasoning. GIF utilizes tagging, structured inputs/outputs, system prompts, instruction-based prompts, and iterative refinement. Indeed, GIF offers a structured alternative to ad hoc prompting, underscoring the need for systematic guidance of LLMs on complex scientific tasks. Using GIF, we evaluate current generalist LLMs' ability to use reasoning towards fragmentation and to perform intensity prediction after fine-tuning. We benchmark performance on a novel QA dataset, the MassSpecGym QA-sim dataset, that we derive from the MassSpecGym dataset. Through these implementations of GIF, we find that GPT-4o and GPT-4o-mini achieve a cosine similarity of 0.36 and 0.35 between the simulated and true spectra, respectively, outperforming other pretrained models including GPT-5, Llama-3.1, and ChemDFM, despite GPT-5's recency and ChemDFM's domain specialization. GIF outperforms several deep learning baselines. Our evaluation of GIF highlights the value of using LLMs not only for spectra simulation but for enabling human-in-the-loop workflows and structured, explainable reasoning in molecular fragmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。