arXiv:2503.20527cs.CLcs.AI2025-03ACL被引 14

用7000多个真实API训练模型,让大模型像镜子一样精准模拟工具响应。

StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs

  • 训练专用大模型模仿7000+真实API的响应行为。
  • 在新构建的MirrorAPI-Bench上表现优于现有方法。
  • 适合需要稳定、真实工具环境的模型评测与训练者使用。

大型语言模型(LLMs)的快速发展推动了工具学习的研究,即通过外接工具使模型完成复杂任务。然而,现有工具环境在稳定性、可扩展性和真实性之间难以平衡,尤其在基准测试中面临挑战。为此,我们提出MirrorAPI框架,通过训练专用大模型来精确模拟真实API的响应,实现对工具环境的“镜像”复制。基于来自7000多个API的请求-响应对数据集,采用监督微调和思维链推理提升仿真保真度。实验表明,MirrorAPI在新构建的MirrorAPI-Bench上的准确率和稳定性均优于当前先进方法,并已集成至StableToolBench中。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has spurred significant interest in tool learning, where LLMs are augmented with external tools to tackle complex tasks. However, existing tool environments face challenges in balancing stability, scalability, and realness, particularly for benchmarking purposes. To address this problem, we propose MirrorAPI, a novel framework that trains specialized LLMs to accurately simulate real API responses, effectively acting as "mirrors" to tool environments. Using a comprehensive dataset of request-response pairs from 7,000+ APIs, we employ supervised fine-tuning and chain-of-thought reasoning to enhance simulation fidelity. MirrorAPI achieves superior accuracy and stability compared to state-of-the-art methods, as demonstrated by its performance on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.

工具学习大模型仿真环境基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。