用偏好和演示数据联合训练奖励模型,提升复杂任务学习效率
Learning from Preferences and Mixed Demonstrations in General Settings
- 提出基于观察的理性偏序奖励框架,统一处理多种人类反馈
- 在少量偏好与示范数据下,性能显著超越现有基线方法
- 适合需要多类型反馈融合的通用强化学习场景
强化学习是序列决策任务中的通用方法,但在复杂任务中设计有效奖励函数往往困难。此时可借助人类偏好或专家示范。然而,现有结合两者的方法常为特定领域设计、缺乏可扩展性。本文提出一种新范式——基于观察的奖励理性偏序(reward-rational partial orderings over observations),具备良好灵活性与可扩展性。基于此,我们提出实用算法 LEOPARD:从偏好与排序示范中学习估计目标。LEOPARD 能高效利用多种数据类型,包括负向示范,在广泛任务中学习奖励函数。实验表明,在有限偏好与示范反馈条件下,其性能显著优于现有基线。进一步研究不同反馈类型的组合效果,发现多类型反馈融合通常更优。
原文摘要 · Abstract (English)
Reinforcement learning is a general method for learning in sequential settings, but it can often be difficult to specify a good reward function when the task is complex. In these cases, preference feedback or expert demonstrations can be used instead. However, existing approaches utilising both together are often ad-hoc, rely on domain-specific properties, or won't scale. We develop a new framing for learning from human data, \emph{reward-rational partial orderings over observations}, designed to be flexible and scalable. Based on this we introduce a practical algorithm, LEOPARD: Learning Estimated Objectives from Preferences And Ranked Demonstrations. LEOPARD can learn from a broad range of data, including negative demonstrations, to efficiently learn reward functions across a wide range of domains. We find that when a limited amount of preference and demonstration feedback is available, LEOPARD outperforms existing baselines by a significant margin. Furthermore, we use LEOPARD to investigate learning from many types of feedback compared to just a single one, and find that combining feedback types is often beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。