arXiv:2509.14480cs.CLcs.AI2025-09被引 12

用语音文本交替训练让智能体学会复杂工具使用。

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

  • 用大模型当裁判,逐轮评估强化学习中的奖励。
  • 在文本基准上任务成功率提升超6%。
  • 适合开发能听会说的交互式智能体。

有效的交互式工具使用需要智能体掌握工具融合推理(TIR):涉及多轮规划与长上下文对话管理的复杂过程。为训练此类动态过程中的多模态智能体,我们引入一个支持语音-文本交替回放的强化学习沙盒环境。核心策略为逐轮仲裁强化学习(TARL),通过大语言模型(LLM)作为裁判提供逐轮评估,解决长时程任务中的信用分配难题。为增强探索能力,我们采用包含数学推理问题的混合任务训练课程。该统一方法使文本基准τ-bench上的任务通过率相比强基线提升超过6%。关键的是,我们证明了该框架适用于微调多模态基础模型以执行代理任务。通过在交错语音-文本回放数据上训练基础多模态LLM,使其具备工具使用能力,为更自然、语音驱动的交互式智能体铺平道路。

原文摘要 · Abstract (English)

Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To train agents for this dynamic process, particularly in multi-modal contexts, we introduce a sandbox environment for reinforcement learning (RL) that supports interleaved speech-text rollouts. Our core strategy, Turn-level Adjudicated Reinforcement Learning (TARL), addresses the challenge of credit assignment in long-horizon tasks by employing a Large Language Model (LLM) as a judge to provide turn-level evaluation. To enhance exploration, we integrate a mixed-task training curriculum with mathematical reasoning problems. This unified approach boosts the task pass rate on the text-based $τ$-bench by over 6% compared to strong RL baselines. Crucially, we demonstrate our framework's suitability for fine-tuning a multi-modal foundation model for agentic tasks. By training a base multi-modal LLM on interleaved speech-text rollouts, we equip it with tool-use abilities, paving the way for more natural, voice-driven interactive agents.

强化学习多模态语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。