arXiv:2507.05458cs.RO2025-07被引 1

通过反事实推理生成更优轨迹,提升机器人偏好学习效率

Counterfactual Reasoning and Environment Design for Active Preference Learning

  • 结合环境设计与反事实推理生成多样化轨迹
  • 在网格世界和真实导航中显著提升奖励学习精度
  • 适合需要高效适应人类偏好的机器人系统

为实现机器人在现实场景中的有效部署,需使其能适应人类偏好,如配送路径中平衡距离、时间与安全。主动偏好学习(APL)通过展示轨迹并让人类排序来学习人类奖励函数。然而现有方法常难以充分探索轨迹空间,且在长时序任务中难以识别有信息量的查询。本文提出CRED,一种用于APL的轨迹生成方法,通过联合优化环境设计与轨迹选择来改进奖励估计。CRED通过环境设计“想象”新场景,并利用反事实推理——从当前信念采样奖励并提问“如果该奖励是真实偏好会怎样?”——生成多样且有信息量的轨迹集供排序。在GridWorld和基于OpenStreetMap数据的真实导航实验中,CRED均显著提升了奖励学习效果,并展现出跨环境的良好泛化能力。

原文摘要 · Abstract (English)

For effective real-world deployment, robots should adapt to human preferences, such as balancing distance, time, and safety in delivery routing. Active preference learning (APL) learns human reward functions by presenting trajectories for ranking. However, existing methods often struggle to explore the full trajectory space and fail to identify informative queries, particularly in long-horizon tasks. We propose CRED, a trajectory generation method for APL that improves reward estimation by jointly optimizing environment design and trajectory selection. CRED "imagines" new scenarios through environment design and uses counterfactual reasoning -- by sampling rewards from its current belief and asking "What if this reward were the true preference?" -- to generate a diverse and informative set of trajectories for ranking. Experiments in GridWorld and real-world navigation using OpenStreetMap data show that CRED improves reward learning and generalizes effectively across different environments.

偏好学习机器人反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。