arXiv:2507.00611cs.LGcs.AI2025-07被引 2

用先验知识提升机器人强化学习的样本效率

Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

  • 将真实奖励拆分为先验项与可学习残差项
  • 在多个环境中使采样效率显著提升,物理机器人实验成功率达90%以上
  • 适合需要快速训练的机器人任务,尤其可用语言或工程经验做先验

基于偏好强化学习(PbRL)为复杂机器人环境提供了替代启发式奖励设计的前景。然而,PbRL常因样本效率低而受限,需大量昂贵的人类反馈。已有工作尝试从示范中学习奖励模型,并用偏好进行微调。但当模型为神经网络时,不同训练阶段间损失函数切换易导致优化不稳定和性能下降。本文提出残差奖励模型(RRM),假设环境真实奖励可分解为先验奖励与学习的残差项之和。先验奖励可来自工程启发、语言生成或逆强化学习获得,学习部分则通过偏好数据训练。在Meta-World和DM-Control上的实验表明,不同先验类型下,RRM显著提升了常见PbRL方法的样本效率。进一步在真实Franka Panda机器人上验证,该方法加速策略学习,以更少步数实现超过90%的成功率。

原文摘要 · Abstract (English)

Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, transitioning between different loss functions across training phases often leads to unstable optimization and performance degradation. In this paper, we propose a method to effectively leverage prior knowledge with a Residual Reward Model (RRM). An RRM assumes that the true reward of the environment can be split into a sum of two parts: a prior reward and a learned reward. The prior reward is a term available before training, such as an engineering heuristic ``best guess'', a language-generated reward, or a reward function learned from inverse reinforcement learning, and the learned reward is then trained with preferences as a residual offset. Experimental results in Meta-World and DM-Control show that RRMs substantially improve the sample efficiency of common PbRL methods across various prior reward types. Furthermore, we demonstrate the practical efficacy of our method on a physical Franka Panda robot, accelerating policy learning and achieving high success rates in fewer steps than baselines.

强化学习机器人偏好学习残差模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。