arXiv:2508.18669cs.AI2025-08被引 27

让AI agent通过与模拟用户多轮对话,学会动态调用工具解决问题。

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use

  • 用大模型模拟真实用户,在强化学习中动态互动训练代理
  • 在多个任务上表现超越或媲美更大模型,最高达82.5分
  • 适合研究智能体交互、工具使用与强化学习融合的开发者

随着智能体智能的快速发展,大语言模型中的智能体工具使用变得愈发重要。在智能体与用户多轮交互过程中,用户需求具有动态性、不确定性和随机性,给智能体的工具调用能力带来重大挑战。智能体不再只需简单调用工具输出结果,还需通过持续沟通迭代理解用户需求,并同步调用工具解决查询。现有强化学习方法在工具使用中缺乏真正动态用户的集成。为此,我们提出MUA-RL(面向智能体工具使用的多轮用户交互强化学习),首次在该领域将大模型模拟用户融入强化学习循环。MUA-RL旨在使模型自主学习高效沟通并灵活调用工具,以应对动态多轮交互中的实际问题。在多个多轮工具使用基准测试上进行评估,MUA-RL-32B在TAU2 Retail上取得67.3分,在TAU2 Airline上得45.4分,在TAU2 Telecom上得28.3分,在BFCL-V3 Multi Turn上得28.4分,在ACEBench Agent上得82.5分,性能优于或匹配更大规模开源模型如DeepSeek-V3-0324和Qwen3-235B-A22B在非思考设置下的表现。

原文摘要 · Abstract (English)

With the recent rapid advancement of Agentic Intelligence, agentic tool use in LLMs has become increasingly important. During multi-turn interactions between agents and users, the dynamic, uncertain, and stochastic nature of user demands poses significant challenges to the agent's tool invocation capabilities. Agents are no longer expected to simply call tools to deliver a result; rather, they must iteratively refine their understanding of user needs through communication while simultaneously invoking tools to resolve user queries. Existing reinforcement learning (RL) approaches for tool use lack the integration of genuinely dynamic users during the RL training process. To bridge this gap, we introduce MUA-RL (Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use), a novel reinforcement learning framework that, for the first time in the field of agentic tool use, integrates LLM-simulated users into the reinforcement learning loop. MUA-RL aims to enable autonomous learning of models to communicate with users efficiently and use various tools to solve practical problems in dynamic multi-turn interactions. Evaluations are done on several multi-turn tool-using benchmarks (see Figure 1). Specifically, MUA-RL-32B achieves 67.3 on TAU2 Retail, 45.4 on TAU2 Airline, 28.3 on TAU2 Telecom, 28.4 on BFCL-V3 Multi Turn, and 82.5 on ACEBench Agent -- outperforming or matching the performance of larger open-source models such as DeepSeek-V3-0324 and Qwen3-235B-A22B in non-thinking settings.

智能体强化学习工具使用多轮交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。