用语言设计分层奖励,让AI更听话地完成复杂任务。
Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications
- 从人类语言中提取分层行为规范,生成多层级奖励信号
- 在长周期任务中显著提升任务完成率与规范遵守度
- 适合需要精细控制行为的智能体应用,如机器人操作
训练人工智能执行任务时,人类不仅关注任务是否完成,还关心执行过程。随着智能体处理的任务日益复杂,使其行为与人类规范对齐成为负责任部署的关键。奖励设计通过将人类期望转化为引导强化学习的奖励函数,提供了直接对齐路径。然而,现有方法难以捕捉长周期任务中的细微人类偏好。为此,本文提出分层奖励设计(HRDL):一种扩展经典奖励设计的框架,用于编码复杂任务中更丰富的行为规范。进一步提出语言转分层奖励(L2HR)作为解决方案。实验表明,采用L2HR设计奖励的智能体不仅能高效完成任务,且更严格遵守人类规范。HRDL与L2HR共同推进了人类对齐型智能体的研究。
原文摘要 · Abstract (English)
When training artificial intelligence (AI) to perform tasks, humans often care not only about whether a task is completed but also how it is performed. As AI agents tackle increasingly complex tasks, aligning their behavior with human-provided specifications becomes critical for responsible AI deployment. Reward design provides a direct channel for such alignment by translating human expectations into reward functions that guide reinforcement learning (RL). However, existing methods are often too limited to capture nuanced human preferences that arise in long-horizon tasks. Hence, we introduce Hierarchical Reward Design from Language (HRDL): a problem formulation that extends classical reward design to encode richer behavioral specifications for hierarchical RL agents. We further propose Language to Hierarchical Rewards (L2HR) as a solution to HRDL. Experiments show that AI agents trained with rewards designed via L2HR not only complete tasks effectively but also better adhere to human specifications. Together, HRDL and L2HR advance the research on human-aligned AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。