构建生物医学工具调用数据集,提升大模型在基因组等领域的专业能力。
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models

- 收集34个生物数据库工具,构建7040条人工验证的查询-调用对。
- 微调后模型性能超越GPT-5.1,专家评估显示答案质量显著提升。
- 适合生物医药领域研究者、AI辅助科研开发者使用。
尽管大语言模型在通用任务上表现优异,但在生物医学等高度专业化领域仍表现不佳。主要瓶颈在于模型难以有效调用生物医学工具,而临床与科研人员日常工作中高度依赖这些工具。现有通用工具调用数据集虽提升了模型能力,但生物医学领域仍依赖上下文学习,且工具种类有限。为此,我们提出BioTool,一个用于微调大模型的综合性生物医学工具调用数据集。该数据集包含来自NCBI、Ensembl和UniProt的34个常用工具,以及7,040条高质量、人工验证的查询-API调用配对,覆盖变异、基因组学、蛋白质组学、进化和一般生物学。在40亿参数模型上微调BioTool后,模型在生物医学工具调用任务中表现显著优于当前领先的商用模型(如GPT-5.1)。人类专家评估进一步表明,集成经BioTool微调的工具调用器后,下游回答质量明显优于未使用工具的同模型。完整数据集与评估代码已开源:https://github.com/gxx27/BioTool。
原文摘要 · Abstract (English)
Despite the success of large language models (LLMs) on general-purpose tasks, their performance in highly specialized domains such as biomedicine remains unsatisfactory. A key limitation is the inability of LLMs to effectively leverage biomedical tools, which clinical experts and biomedical researchers rely on extensively in daily workflows. While recent general-domain tool-calling datasets have substantially improved the capabilities of LLM agents, existing efforts in the biomedical domain largely rely on in-context learning and restrict models to a small set of tools. To address this gap, we introduce BioTool, a comprehensive biomedical tool-calling dataset designed for fine-tuning LLMs. BioTool comprises 34 frequently used tools collected from the NCBI, Ensembl, and UniProt databases, along with 7,040 high-quality, human-verified query-API call pairs spanning variation, genomics, proteomics, evolution, and general biology. Fine-tuning a 4-billion-parameter LLM on BioTool yields substantial improvements in biomedical tool-calling performance, outperforming cutting-edge commercial LLMs such as GPT-5.1. Furthermore, human expert evaluations demonstrate that integrating a BioTool-fine-tuned tool caller significantly improves downstream answer quality compared to the same LLM without tool usage, highlighting the effectiveness of BioTool in enhancing the biomedical capabilities of LLMs. The full dataset and evaluation code are available at https://github.com/gxx27/BioTool
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。