arXiv:2509.22402cs.LGcs.RO2025-09被引 1

用视觉关键点预测距离,自动生成机器人操作的密集奖励。

ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation

  • 从无动作视频中提取关键点,隐式推断空间距离。
  • 构建可预测中间目标的规划模型,实现结构化奖励设计。
  • 在长程操作任务中显著加速学习,优于现有方法。

视觉强化学习中的奖励设计仍是机器人操作的关键瓶颈。模拟环境中通常基于与目标位置的距离设计奖励,但真实视觉场景因感知限制难以获取精确位置信息。本文提出一种通过图像关键点隐式推断空间距离的方法,并引入奖励学习中的预期模型(ReLAM),可从无动作视频示范中自动生成密集、结构化的奖励信号。ReLAM首先学习一个作为规划器的预期模型,在最优路径上提出基于关键点的中间子目标,构建与任务几何目标直接对齐的结构化学习课程。基于这些预期子目标,提供连续奖励信号,在分层强化学习框架下训练低层目标条件策略,具备可证明的次优性界。在复杂、长时程操作任务上的大量实验表明,ReLAM显著加速学习并取得优于当前先进方法的性能。

原文摘要 · Abstract (English)

Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on the distance to a target position. However, such precise positional information is often unavailable in real-world visual settings due to sensory and perceptual limitations. In this study, we propose a method that implicitly infers spatial distances through keypoints extracted from images. Building on this, we introduce Reward Learning with Anticipation Model (ReLAM), a novel framework that automatically generates dense, structured rewards from action-free video demonstrations. ReLAM first learns an anticipation model that serves as a planner and proposes intermediate keypoint-based subgoals on the optimal path to the final goal, creating a structured learning curriculum directly aligned with the task's geometric objectives. Based on the anticipated subgoals, a continuous reward signal is provided to train a low-level, goal-conditioned policy under the hierarchical reinforcement learning (HRL) framework with provable sub-optimality bound. Extensive experiments on complex, long-horizon manipulation tasks show that ReLAM significantly accelerates learning and achieves superior performance compared to state-of-the-art methods.

机器人操作强化学习视觉感知奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。