提出可量化奖励与人类偏好对齐程度的新指标,提升强化学习奖励设计效率。
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
- 引入轨迹对齐系数衡量奖励函数与人类排序的一致性。
- 在11人实验中,使用该指标使奖励选择成功率提升41%。
- 适合需要优化奖励设计的RL实践者,降低认知负担。
强化学习代理的根本限制在于其学习的奖励函数质量,但奖励设计常被忽视,假设良好定义的奖励易于获得。然而实践中,设计奖励困难,且评估其正确性同样棘手:如何判断奖励函数是否正确?本文聚焦奖励对齐问题——评估奖励函数是否准确反映人类利益相关者的偏好。我们提出轨迹对齐系数(Trajectory Alignment Coefficient),量化人类对轨迹分布排序与奖励函数诱导出的排序之间的相似性。该指标具备理想性质:无需真实奖励、对基于势能的奖励重塑不变、适用于在线强化学习。在11名强化学习实践者的用户研究中,使用该指标进行奖励选择显著提升表现:相比仅依赖奖励函数,认知负荷降低1.5倍,82%用户更偏好该工具,奖励选择成功率达41%提升。
原文摘要 · Abstract (English)
Reinforcement learning agents are fundamentally limited by the quality of the reward functions they learn from, yet reward design is often overlooked under the assumption that a well-defined reward is readily available. However, in practice, designing rewards is difficult, and even when specified, evaluating their correctness is equally problematic: how do we know if a reward function is correctly specified? In our work, we address these challenges by focusing on reward alignment -- assessing whether a reward function accurately encodes the preferences of a human stakeholder. As a concrete measure of reward alignment, we introduce the Trajectory Alignment Coefficient to quantify the similarity between a human stakeholder's ranking of trajectory distributions and those induced by a given reward function. We show that the Trajectory Alignment Coefficient exhibits desirable properties, such as not requiring access to a ground truth reward, invariance to potential-based reward shaping, and applicability to online RL. Additionally, in an 11 -- person user study of RL practitioners, we found that access to the Trajectory Alignment Coefficient during reward selection led to statistically significant improvements. Compared to relying only on reward functions, our metric reduced cognitive workload by 1.5x, was preferred by 82% of users and increased the success rate of selecting reward functions that produced performant policies by 41%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。