arXiv:2504.01400cs.CLcs.AI2025-04AAAI被引 5

让大模型学会自动优化调用工具,提升复杂任务解决能力。

ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning

  • 根据模型能力动态调整训练数据,持续激发潜力
  • 无需外部反馈,模型可自我迭代优化工具调用
  • 测试时自适应决定停止时机,高效且省资源

工具学习使大语言模型能够调用外部工具以解决复杂用户任务,成为扩展模型能力的有力途径。然而,现有方法多聚焦于微调数据合成,忽视了充分激发模型潜能。本文提出ToolACE-R框架,包含模型感知的迭代训练与自适应精炼机制。该框架通过动态调整训练样本,依据模型演进能力最大化其潜力;引入自精炼训练语料,强化模型迭代优化工具调用的能力,无需外部反馈即可提升性能;并设计自适应自精炼机制,使模型在测试时可自主判断何时停止迭代,实现高效推理。在多个基准数据集上的实验表明,ToolACE-R性能媲美先进API模型,且通过自适应精炼可进一步提升工具调用效果。结果验证了其有效性和通用性,为更高效、可扩展的工具学习提供了新方向。

原文摘要 · Abstract (English)

Tool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model's evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement mechanism for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced API-based models. The performance of tool invocation can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning.

工具学习大模型自迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。