自进化数据+可验证奖励,让智能体高效学会多轮工具使用。
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
- 用自迭代生成数据与可验证奖励机制,解决多轮工具交互训练难题。
- 在tau^2-bench上达73.0%(航空)和98.3%(电信)通过率,超越现有模型。
- 适合研究复杂对话系统、强化学习与自动化工具应用的开发者。
交互式工具使用智能体需通过与人类及外部环境的多轮互动完成真实任务,涉及对话状态追踪、多步工具执行及复杂指令遵循。后训练此类智能体极具挑战:高质量多轮工具使用数据难以规模化合成,而强化学习可能因用户模拟产生噪声信号,降低训练效率。本文提出统一框架,结合自进化数据代理与基于验证器的强化学习。系统EigenData为分层多智能体引擎,可合成带工具依据的对话并生成每实例可执行的校验器,通过闭环自进化过程更新提示与流程,提升生成可靠性。基于合成数据,设计一种强化学习方案:先微调用户模型,再采用类似GRPO的轨迹级组相对优势与动态过滤,实现持续优化。在tau^2-bench上,最佳模型在航空任务中达到73.0% pass^1,在电信任务中达98.3% pass^1,性能匹配或超越前沿模型。结果表明,该方法无需昂贵人工标注即可规模化构建复杂工具行为。
原文摘要 · Abstract (English)
Interactive tool-using agents must solve real-world tasks via multi-turn interaction with both humans and external environments, requiring dialogue state tracking, multi-step tool execution, while following complex instructions. Post-training such agents is challenging because synthesis for high-quality multi-turn tool-use data is difficult to scale, and reinforcement learning (RL) could face noisy signals caused by user simulation, leading to degraded training efficiency. We propose a unified framework that combines a self-evolving data agent with verifier-based RL. Our system, EigenData, is a hierarchical multi-agent engine that synthesizes tool-grounded dialogues together with executable per-instance checkers, and improves generation reliability via closed-loop self-evolving process that updates prompts and workflow. Building on the synthetic data, we develop an RL recipe that first fine-tunes the user model and then applies GRPO-style training with trajectory-level group-relative advantages and dynamic filtering, yielding consistent improvements beyond SFT. Evaluated on tau^2-bench, our best model reaches 73.0% pass^1 on Airline and 98.3% pass^1 on Telecom, matching or exceeding frontier models. Overall, our results suggest a scalable pathway for bootstrapping complex tool-using behaviors without expensive human annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。