arXiv:2602.11393cs.RO2026-02

用视觉运动预测建模人类偏好,提升机器人从第一视角视频学技能的效果

Human Preference Modeling Using Visual Motion Prediction Improves Robot Skill Learning from Egocentric Human Video

  • 通过预测图像间关键点运动来建模人类偏好
  • 在真实机器人上仅用10次演示即实现性能超越
  • 适合需要从人视频学习复杂动作的机器人研究者

本文提出一种基于第一视角人类视频的机器人技能学习方法,通过建模人类偏好作为奖励函数来优化机器人行为。以往方法依赖视觉状态与示范终点之间的时序距离来衡量长期价值,但受限于假设且难以跨实体和环境迁移。本方法通过学习连续图像间追踪点的运动预测,将预测与实际观察到的物体运动的一致性定义为每一步的奖励信号。采用改进的Soft Actor Critic(SAC)算法,以10次真实机器人示范初始化,直接在机器人上估计值函数并优化策略。实验表明,该方法可在仿真和真实机器人上多个任务中达到或优于先前工作表现。

原文摘要 · Abstract (English)

We present an approach to robot learning from egocentric human videos by modeling human preferences in a reward function and optimizing robot behavior to maximize this reward. Prior work on reward learning from human videos attempts to measure the long-term value of a visual state as the temporal distance between it and the terminal state in a demonstration video. These approaches make assumptions that limit performance when learning from video. They must also transfer the learned value function across the embodiment and environment gap. Our method models human preferences by learning to predict the motion of tracked points between subsequent images and defines a reward function as the agreement between predicted and observed object motion in a robot's behavior at each step. We then use a modified Soft Actor Critic (SAC) algorithm initialized with 10 on-robot demonstrations to estimate a value function from this reward and optimize a policy that maximizes this value function, all on the robot. Our approach is capable of learning on a real robot, and we show that policies learned with our reward model match or outperform prior work across multiple tasks in both simulation and on the real robot.

机器人学习视觉运动偏好建模第一人称视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。