开源150万条真实工具调用数据,提升大模型任务自动化能力
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- 基于真实MCP环境生成多样工具调用轨迹
- 在BFCL V3上超越更大闭源模型表现
- 适合研究智能体、工具使用与自动化系统者
大型语言模型智能体正迅速成为跨领域自动化任务的强大工具。然而,开源社区的发展受限于高质量、许可宽松的工具型智能体训练数据的缺乏。现有数据集在多样性、真实性和复杂性方面存在局限,尤其在多工具、多轮交互方面。为此,我们提出Toucan,目前最大的公开可用工具型智能体数据集,包含从近500个真实世界模型上下文协议(MCPs)中合成的150万条轨迹。不同于以往工作,Toucan利用真实的MCP环境生成多样化、真实且具有挑战性的任务,包含真实工具执行的轨迹。我们的流水线首先使用五种不同模型生成广泛的工具使用查询,经基于模型的质量过滤后,再通过三种教师模型和两种智能体框架生成智能体轨迹。严格的规则与模型验证确保输出质量。我们还引入三种扩展机制,进一步丰富任务并模拟多轮对话。在Toucan上微调的模型,在BFCL V3基准上表现优于更大的闭源模型,并在MCP-Universe Bench上推动了帕累托前沿。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents are rapidly emerging as powerful systems for automating tasks across domains. Yet progress in the open-source community is constrained by the lack of high quality permissively licensed tool-agentic training data. Existing datasets are often limited in diversity, realism, and complexity, particularly regarding multi-tool and multi-turn interactions. To address this gap, we introduce Toucan, the largest publicly available tool-agentic dataset to date, containing 1.5 million trajectories synthesized from nearly 500 real-world Model Context Protocols (MCPs). Unlike prior work, Toucan leverages authentic MCP environments to generate diverse, realistic, and challenging tasks with trajectories involving real tool execution. Our pipeline first produces a broad spectrum of tool-use queries using five distinct models, applies model-based quality filtering, and then generates agentic trajectories with three teacher models using two agentic frameworks. Rigorous rule-based and model-based validation ensures high-quality outputs. We also introduce three extension mechanisms to further diversify tasks and simulate multi-turn conversations. Models fine-tuned on Toucan outperform larger closed-source counterparts on the BFCL V3 benchmark and push the Pareto frontier forward on MCP-Universe Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。