用相对反馈提升强化学习采样效率,支持环境与偏好变化适应。
CueLearner: Bootstrapping and local policy adaptation from relative feedback
- 通过相对反馈引导探索,结合离策略RL实现高效学习。
- 在两个稀疏奖励任务中显著提升样本效率,加速收敛。
- 适用于真实导航场景,可动态适应环境或用户偏好变化。
人类指导已成为增强强化学习(RL)的重要工具。然而,传统指导形式如示范或二元标量反馈在收集上存在困难或信息量低,促使研究者探索其他形式的人类输入。相对反馈(如“再往左一点”)在易用性与信息丰富性之间取得了良好平衡。已有研究证明相对反馈可用于改进策略搜索方法,但以往工作局限于特定策略类别,且反馈利用效率不高。本文提出一种新方法,从相对反馈中学习并结合离策略强化学习。在两个稀疏奖励任务上的评估表明,该方法能有效引导探索过程,提升样本效率。此外,方法可使策略适应环境变化或用户偏好调整。最后,我们在真实世界导航任务中验证了其可行性,在稀疏奖励设置下成功学习导航策略。
原文摘要 · Abstract (English)
Human guidance has emerged as a powerful tool for enhancing reinforcement learning (RL). However, conventional forms of guidance such as demonstrations or binary scalar feedback can be challenging to collect or have low information content, motivating the exploration of other forms of human input. Among these, relative feedback (i.e., feedback on how to improve an action, such as "more to the left") offers a good balance between usability and information richness. Previous research has shown that relative feedback can be used to enhance policy search methods. However, these efforts have been limited to specific policy classes and use feedback inefficiently. In this work, we introduce a novel method to learn from relative feedback and combine it with off-policy reinforcement learning. Through evaluations on two sparse-reward tasks, we demonstrate our method can be used to improve the sample efficiency of reinforcement learning by guiding its exploration process. Additionally, we show it can adapt a policy to changes in the environment or the user's preferences. Finally, we demonstrate real-world applicability by employing our approach to learn a navigation policy in a sparse reward setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。