arXiv:2510.14944cs.CLcs.AI2025-10ACL被引 1

首个代谢组学大模型评测基准,揭示模型在跨数据库匹配上的短板。

MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics

  • 构建多任务评测框架,覆盖知识、理解、定位、推理与科研五大能力。
  • 25个模型测试显示,长尾代谢物标注稀疏时性能显著下降。
  • 适合生物信息学与AI交叉研究者,推动精准代谢分析工具发展。

大语言模型在通用文本任务中表现优异,但在需深度关联知识的科学领域能力仍不明确。代谢组学面临生化通路复杂、标识系统异构、数据库碎片化等挑战。为此,我们提出首个代谢组学评估基准MetaBench,基于权威公开资源构建,涵盖知识、理解、定位、推理和科研五项核心能力。对25个开源与闭源模型的评估发现:尽管文本生成表现良好,跨数据库标识符定位仍具挑战性,尤其在长尾代谢物(注释稀疏)上性能明显下降。MetaBench为开发与评估代谢组学AI系统提供关键基础设施,助力实现可靠计算工具的系统性进步。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics presents unique challenges with its complex biochemical pathways, heterogeneous identifier systems, and fragmented databases. To systematically evaluate LLM capabilities in this domain, we introduce MetaBench, the first benchmark for metabolomics assessment. Curated from authoritative public resources, MetaBench evaluates five capabilities essential for metabolomics research: knowledge, understanding, grounding, reasoning, and research. Our evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks: while models perform well on text generation tasks, cross-database identifier grounding remains challenging even with retrieval augmentation. Model performance also decreases on long-tail metabolites with sparse annotations. With MetaBench, we provide essential infrastructure for developing and evaluating metabolomics AI systems, enabling systematic progress toward reliable computational tools for metabolomics research.

代谢组学大模型评测生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。