让大模型代理更主动、个性化,像同事一样协作
Training Proactive and Personalized LLM Agents
- 提出PPP三维度:效率、主动性、个性化,指导代理训练
- 在真实任务中超越GPT-5等基线平均16.7分,提问更精准
- 用用户模拟器和反馈优化代理,适合需要长期协作的场景
尽管进展迅速,当前AI代理主要针对孤立任务优化。我们主张转向以协作为核心的训练范式,使代理能与人沟通并适应个体需求。为推动这一转变,我们正式定义协作型代理的三个维度:生产力、主动性与个性化(PPP)。提出UserVille交互环境,包含可配置的基于LLM的用户模拟器和以用户为中心的反馈机制,用于评估这些维度。设计多目标强化学习框架,通过任务完成奖励、问题代价和偏好一致性奖励联合优化。在两个真实世界代理任务(SWE-Bench 和 BrowseComp-Plus)上,经过PPP训练的代理平均比强基线(包括GPT-5)高出16.7分,提问更聚焦,并能泛化到未见偏好与任务。后续用户研究进一步验证了以用户为中心反馈对训练高效且易监督协作代理的重要性。
原文摘要 · Abstract (English)
Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift in real-world complex applications, we first formalize three dimensions of collaborative AI agents: Productivity, Proactivity, and Personalization (PPP). We introduce UserVille, an interactive environment with configurable LLM-based user simulators and user-centric feedback to evaluate these dimensions, and propose a multi-objective reinforcement learning framework that optimizes them using rewards from task outcomes, question effort, and preference adherence. On two real-world agentic tasks (SWE-Bench and BrowseComp-Plus), PPP-trained agents outperform strong LLM baselines (including GPT-5) by an average of 16.7 points, ask more targeted questions, and generalize to unseen preferences and tasks. A follow-up user study further highlights the importance of user-centric feedback for training collaborative agents that are both more effective and easier to supervise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。