arXiv:2411.17764cs.ROcs.AI2024-11ICCV被引 10

用视频自监督学习任务进度,让机器人无监督学会复杂动作。

PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online Refinement

  • 从视频中自监督学习任务进度分布作为奖励信号。
  • 在线训练时对抗性修正异常观测,缓解分布偏移问题。
  • 无需人工标注,适合真实场景下机器人自主学习。

我们提出 PROGRESSOR,一种从视频中学习任务无关奖励函数的新框架,使策略可通过目标条件强化学习(RL)在无手动标注的情况下进行训练。其核心是通过当前、初始和目标观测学习任务进度的分布估计,并以自监督方式优化。关键在于,在线强化学习过程中,通过对抗性回推机制对分布外观测的预测进行修正,以缓解非专家观测带来的分布偏移问题。结合进度预测的密集奖励与对抗性回推,PROGRESSOR 实现了无需外部监督的复杂行为学习。该模型在 EPIC-KITCHENS 的大规模第一人称人类视频上预训练,无需领域特定数据微调即可在嘈杂示范下实现真实机器人离线强化学习,性能优于现有提供密集视觉奖励的方法。结果表明,PROGRESSOR 在缺乏直接动作标签和任务特定奖励的场景中具有规模化应用潜力。

原文摘要 · Abstract (English)

We present PROGRESSOR, a novel framework that learns a task-agnostic reward function from videos, enabling policy training through goal-conditioned reinforcement learning (RL) without manual supervision. Underlying this reward is an estimate of the distribution over task progress as a function of the current, initial, and goal observations that is learned in a self-supervised fashion. Crucially, PROGRESSOR refines rewards adversarially during online RL training by pushing back predictions for out-of-distribution observations, to mitigate distribution shift inherent in non-expert observations. Utilizing this progress prediction as a dense reward together with an adversarial push-back, we show that PROGRESSOR enables robots to learn complex behaviors without any external supervision. Pretrained on large-scale egocentric human video from EPIC-KITCHENS, PROGRESSOR requires no fine-tuning on in-domain task-specific data for generalization to real-robot offline RL under noisy demonstrations, outperforming contemporary methods that provide dense visual reward for robotic learning. Our findings highlight the potential of PROGRESSOR for scalable robotic applications where direct action labels and task-specific rewards are not readily available.

强化学习自监督机器人视觉奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。