用单次成功示范生成可靠进度奖励,提升机器人长任务成功率
RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

- 基于参考片段对比学习,从一次成功演示生成密集进度奖励
- 在9个仿真和4个真实任务中表现最佳,长任务如折布提升显著
- 自动过滤不可信匹配,避免错误奖励,适合复杂操作场景
机器人操纵的强化学习常受奖励设计瓶颈制约,尤其在长时序任务中:稀疏的成功奖励提供弱监督,而手工设计的稠密奖励耗时且泛化差。基于进度的奖励模型通过估计观测值向任务完成的推进程度提供替代方案,但现有方法通常需要任务特定演示或进度标签,且可能为视觉合理但物理错误的状态赋予高奖励。我们提出参考锚定奖励模型(RARM),一种轻量级视觉比较器,可将单一成功演示转化为稠密、进度感知的奖励。RARM 在通用视频上以对比时间目标一次性训练,无需机器人特定数据、任务特定奖励标签或每任务奖励工程。部署时,RARM 将回放片段与参考片段匹配,并仅对高置信度前向进展给予奖励,抑制可能产生虚假正向奖励的不确定匹配。在来自 LIBERO 和 MetaWorld 的 9 个仿真操纵任务及 4 个真实世界任务中,RARM 在后续强化学习训练中取得最佳总体成功率,尤其在长时序任务如布料折叠中表现突出,此时不可靠的进度估计危害更大。
原文摘要 · Abstract (English)
Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poorly across tasks. Progress-based reward models offer a promising alternative by estimating how far an observation has advanced toward task completion, but existing approaches often require task-specific demonstrations or progress labels, and can assign high rewards to visually plausible but physically incorrect states. We introduce the Reference-Anchored Reward Model (RARM), a lightweight visual comparator that converts a single successful demonstration into a dense, progress-aware reward. RARM is trained once on general-purpose videos with a contrastive temporal objective, requiring no robot-specific data, task-specific reward labels, or per-task reward engineering. At deployment, RARM matches rollout clips to reference clips and rewards only confident forward progress, suppressing uncertain matches that may otherwise produce false-positive rewards. Across 9 simulated manipulation tasks from LIBERO and MetaWorld and 4 real-world tasks, RARM achieves the best overall success rates in subsequent RL training, with particularly large gains on long-horizon tasks such as cloth folding, where unreliable progress estimates are especially harmful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。