用大模型让普通人也能轻松查询代谢组学知识图谱。
MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs
- 设计多智能体系统,把自然语言转为可执行的查询语句。
- 在真实数据集上准确回答植物与代谢物关系等复杂问题。
- 适合无编程基础的生物研究人员快速挖掘数据价值。
基于质谱的代谢组学产生复杂高维数据,蕴含巨大生物学发现潜力,但难以整合与解读。知识图谱(KGs)通过将谱图、注释、物种、化学类别和生物活性统一为可互操作的网络,整合异构信息;然而其实际应用受限于专用表示与查询语言的学习门槛。本文提出 MetaboT,一个开源的基于大语言模型(LLM)的多智能体框架,可将自然语言问题转化为对代谢组学知识图谱的可执行 SPARQL 查询。MetaboT 通过模块化架构,由专门智能体分别处理范围验证、实体消歧、模式感知查询生成、迭代优化和结果解释,克服单模型方法的幻觉与模式合规性缺陷。我们在实验天然产物知识图谱(ENPKG)上验证 MetaboT,采用专家编写的自然语言问题与参考 SPARQL 查询组成的基准,展示其对植物-代谢物关系及生物活性等复杂问题的回答能力。MetaboT 降低了代谢组学研究者的技术门槛,实现无需专业编程即可进行语义数据挖掘。
原文摘要 · Abstract (English)
Mass spectrometry-based metabolomics generates complex, high-dimensional data that holds vast potential for biological discovery but remains difficult to integrate and interpret. Knowledge graphs (KGs) unify this heterogeneous information by representing spectra, annotations, taxa, chemical classes, and biological activities as a single interoperable network; however, their practical use is limited by the steep learning curve of corresponding specialized representation and query languages. Here we introduce MetaboT, an open-source multi-agent Large Language Model (LLM) framework that translates natural-language questions into executable SPARQL queries over metabolomics knowledge graphs. MetaboT mitigates the hallucination and schema-compliance limitations of single-model approaches through a modular architecture in which specialised agents handle scope validation, entity resolution against authoritative resources, schema-aware query generation, iterative refinement, and result interpretation. We validated MetaboT on the Experimental Natural Products Knowledge Graph (ENPKG), using an expert-authored benchmark of natural-language questions paired with reference SPARQL queries, and demonstrate its ability to answer complex questions about plant--metabolite relationships and biological activities. MetaboT lowers the technical barrier for metabolomics researchers and enables semantic data mining without specialised programming expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。