通过环境设计与反事实推理,让机器人更高效地学习人类偏好。
CRED: Counterfactual Reasoning and Environment Design for Active Preference Learning
- 结合环境设计与反事实推理生成对比轨迹对。
- 在相同样本数下,奖励函数推断准确率显著提升。
- 适合需要高效人机偏好对齐的复杂任务场景。
随着机器人操作环境和任务复杂度提升,显式指定并平衡优化目标以实现期望行为变得愈发困难。此类系统若能根据人类偏好调整行为并响应修正,则性能显著提升,但手动编码反馈不可行。主动偏好学习(APL)通过展示轨迹供用户排序来学习人类奖励函数。然而现有方法依赖固定轨迹集或回放缓冲区,查询多样性不足,常无法识别有信息量的比较。本文提出CRED,一种新型轨迹生成方法,通过联合优化环境设计与轨迹选择,高效地从用户处提取偏好。CRED通过环境设计‘想象’新场景,并利用反事实推理——从当前信念中采样可能的奖励,提出‘如果这是真实偏好会怎样?’的问题——生成能凸显不同奖励函数差异的轨迹对。大量实验与用户研究显示,CRED在奖励准确性与样本效率方面显著优于现有最优方法,且获得更高用户评分。
原文摘要 · Abstract (English)
As a robot's operational environment and tasks to perform within it grow in complexity, the explicit specification and balancing of optimization objectives to achieve a preferred behavior profile moves increasingly farther out of reach. These systems benefit strongly by being able to align their behavior to reflect human preferences and respond to corrections, but manually encoding this feedback is infeasible. Active preference learning (APL) learns human reward functions by presenting trajectories for ranking. However, existing methods sample from fixed trajectory sets or replay buffers that limit query diversity and often fail to identify informative comparisons. We propose CRED, a novel trajectory generation method for APL that improves reward inference by jointly optimizing environment design and trajectory selection to efficiently query and extract preferences from users. CRED "imagines" new scenarios through environment design and leverages counterfactual reasoning -- by sampling possible rewards from its current belief and asking "What if this were the true preference?" -- to generate trajectory pairs that expose differences between competing reward functions. Comprehensive experiments and a user study show that CRED significantly outperforms state-of-the-art methods in reward accuracy and sample efficiency and receives higher user ratings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。