arXiv:2602.11351cs.AIcs.LG2026-02被引 1

让AI更主动省力地完成任务,同时减少用户互动负担。

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

  • 用行为优化框架提升AI多轮交互中的信息获取能力
  • 在任务成功率和用户投入上同时优于现有方法
  • 适合需要高效人机协作的复杂场景应用

主动式大语言模型(LLM)代理旨在多轮交互中主动规划、提问与互动,实现超越被动指令响应的高效任务完成,对真实世界以用户为中心的应用至关重要。近年来,代理强化学习(Agentic RL)成为训练此类代理的有前景方案,使其能在多轮设置中学习长期决策策略。然而,现有流程面临关键挑战:难以平衡任务表现与用户参与度——被动代理无法有效适应用户意图,而过度依赖人工反馈又增加用户负担,形成二者间的帕累托前沿。为突破这一边界,我们提出行为代理优化(BAO),一种代理强化学习框架,通过增强并正则化轮间行为,提升信息获取能力,抑制低效或冗余的用户交互。我们在UserRL基准套件的多个任务上评估BAO,结果表明其在任务性能和用户努力方面均显著优于现有主动代理强化学习基线,且达到甚至超过商业级LLM代理的表现,凸显其在复杂多轮场景中训练主动、以用户为中心的LLM代理的有效性。

原文摘要 · Abstract (English)

Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications. Agentic reinforcement learning (RL) has recently emerged as a promising solution for training such agents in multi-turn settings, allowing them to learn long-horizon decision-making strategies. However, existing pipelines face a critical challenge in balancing task performance with user engagement, as passive agents cannot efficiently adapt to users' intentions while overuse of human feedback increases the burden on users, which forms a Pareto Frontier between these two objectives. To push forward this frontier, we propose Behavior Agentic Optimization (BAO), an agentic RL framework that enhances and regularizes inter-turn behaviors to improve information-gathering capabilities and suppress inefficient or redundant interactions with users. We evaluate BAO on multiple tasks from the UserRL benchmark suite and demonstrate that it substantially outperforms proactive agentic RL baselines in terms of both higher task performance and lower user efforts, while achieving comparable or even superior performance to commercial LLM agents, highlighting its effectiveness for training proactive, user-centric LLM agents in complex multi-turn scenarios. Our website: https://proactive-agentic-rl.github.io/.

智能代理强化学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。