arXiv:2510.13709cs.AIcs.LG2025-10被引 4

用人类能动性最大化训练智能助手,更懂何时该放手。

Training LLM Agents to Empower Humans

  • 以人类能动性为目标优化助手行为,鼓励适时退让
  • 用户实验中助手接受率提升31%,建议减少38%
  • 仅需离线文本数据,无需额外人工反馈

助手机器人不应独自完成任务,而应在关键决策时让渡控制权。现有方法如模仿专家或基于推断奖励的强化学习,常导致代理过度干预。我们提出基于最大化人类能动性的新方法Empower,仅需离线文本数据,实现自监督微调。18人用户研究显示,参与者78%偏好我们的助手(p=0.015),接受率提高31%,建议减少38%。我们还构建了模拟人类程序员的多轮代码辅助环境,结果显示,使用Empower训练的代理使模拟程序员在难题上的成功率平均提升192%(相比SFT基线)。该方法仅依赖离线数据,无需人工反馈或可验证奖励,可规模化实现对齐的有用智能体。

原文摘要 · Abstract (English)

Assistive agents should not only take actions on behalf of a human, but also step out of the way and cede control when there are important decisions to be made. However, current methods for building assistive agents, whether via mimicking expert humans or via RL finetuning on an inferred reward, often encourage agents to complete tasks on their own rather than truly assisting the human attain their objectives. Additionally, these methods often require costly explicit human feedback to provide a training signal. We propose a new approach to tuning assistive language models based on maximizing the human's empowerment, their ability to effect desired changes in the environment. Our empowerment-maximizing method, Empower, only requires offline text data, providing a self-supervised method for fine-tuning language models to better assist humans. To study the efficacy of our approach, we conducted an 18-person user study comparing our empowerment assistant with a strong baseline. Participants preferred our assistant 78% of the time (p=0.015), with a 31% higher acceptance rate and 38% fewer suggestions. Additionally, we introduce a new environment for evaluating multi-turn code assistance using simulated humans. Using this environment, we show that agents trained with Empower increase the success rate of a simulated human programmer on challenging coding questions by an average of 192% over an SFT baseline. With this empowerment objective, we provide a framework for useful aligned AI agents at scale using only offline data without the need for any additional human feedback or verifiable rewards.

AI助手能动性自监督人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。