arXiv:2409.07268cs.LG2024-09ICRA被引 3

让强化学习同时理解人类对行为的‘一样好’和‘更优’反馈,提升教学效率。

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

  • 设计等效偏好学习任务,让模型在行为被标为相同时预测相似奖励。
  • 在10个任务上验证,同时学习等效与明确偏好可提升反馈利用效率。
  • 适合需要高效人类反馈的机器人控制场景,尤其当教师常给出‘差不多’评价时。

基于偏好的强化学习(PBRL)通过人类教师对智能体行为的偏好反馈直接学习,无需精心设计奖励函数。然而现有方法主要依赖显式偏好,忽视了教师可能给出的等效偏好(即两个行为被认为同样好)。这种忽略可能使智能体无法全面理解教师的任务意图,导致信息损失。为此,本文提出等效偏好学习任务,通过促使神经网络在两个行为被标记为等效时输出相似奖励预测来优化模型。在此基础上,提出多类型偏好学习(MTPL)方法,可同时从等效偏好和显式偏好中学习。我们在DeepMind Control Suite的10个运动与机器人操作任务上,将MTPL应用于四个先进基线方法进行验证。实验表明,同时利用等效与显式偏好,能更全面地理解教师反馈,显著提升反馈效率。项目页面: https://github.com/FeiCuiLengMMbb/paper_MTPL

原文摘要 · Abstract (English)

Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teachers may choose equal preferences. This neglect may hinder the understanding of the agent regarding the task perspective of the teacher, leading to the loss of important information. To address this issue, we introduce the Equal Preference Learning Task, which optimizes the neural network by promoting similar reward predictions when the behaviors of two agents are labeled as equal preferences. Building on this task, we propose a novel PBRL method, Multi-Type Preference Learning (MTPL), which allows simultaneous learning from equal preferences while leveraging existing methods for learning from explicit preferences. To validate our approach, we design experiments applying MTPL to four existing state-of-the-art baselines across ten locomotion and robotic manipulation tasks in the DeepMind Control Suite. The experimental results indicate that simultaneous learning from both equal and explicit preferences enables the PBRL method to more comprehensively understand the feedback from teachers, thereby enhancing feedback efficiency. Project page: \url{https://github.com/FeiCuiLengMMbb/paper_MTPL}

强化学习人类偏好机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。