arXiv:2412.07990cs.ROcs.AI2024-12被引 1

通过自适应选择反馈形式,让机器人更高效地学习用户偏好并避开危险行为。

Adaptive Querying for Reward Learning from Human Feedback

  • 先选关键状态,再根据信息量和反馈成本选最优反馈方式。
  • 实验显示该方法比固定反馈方式少用40%的交互样本。
  • 适合需要快速适配用户需求的智能机器人系统。

从人类反馈中学习是训练机器人适应用户偏好并提升安全性的常用方法。现有方法通常采用单一交互形式获取反馈,未充分利用用户与机器人之间的多种互动方式。本文研究如何通过优化查询状态和反馈形式,利用多种人类反馈来学习与不安全行为相关的惩罚函数。提出的自适应反馈选择方法为迭代式两阶段流程:首先选择关键状态进行查询,然后基于信息增益在采样关键状态上选择最优反馈形式。反馈形式的选择还考虑了不同反馈方式的成本与获取概率。仿真实验表明该方法在学习规避不良行为方面具有更高的样本效率。用户实机实验结果进一步验证了该方法在获取有信息量且符合用户意图的反馈方面的实用性与有效性。实验视频、代码及附录见:https://tinyurl.com/AFS-learning。

原文摘要 · Abstract (English)

Learning from human feedback is a popular approach to train robots to adapt to user preferences and improve safety. Existing approaches typically consider a single querying (interaction) format when seeking human feedback and do not leverage multiple modes of user interaction with a robot. We examine how to learn a penalty function associated with unsafe behaviors using multiple forms of human feedback, by optimizing both the query state and feedback format. Our proposed adaptive feedback selection is an iterative, two-phase approach which first selects critical states for querying, and then uses information gain to select a feedback format for querying across the sampled critical states. The feedback format selection also accounts for the cost and probability of receiving feedback in a certain format. Our experiments in simulation demonstrate the sample efficiency of our approach in learning to avoid undesirable behaviors. The results of our user study with a physical robot highlight the practicality and effectiveness of adaptive feedback selection in seeking informative, user-aligned feedback that accelerate learning. Experiment videos, code and appendices are found on our website: https://tinyurl.com/AFS-learning.

强化学习人机交互自适应反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。