arXiv:2502.20630cs.ROcs.AI2025-02ICLR被引 14

用分段演示视频自动学奖励,让机器人更懂任务细节。

Subtask-Aware Visual Reward Learning from Segmented Demonstrations

  • 将视频按子任务分割,当作真实奖励信号训练模型
  • 在Meta-World和家具装配任务中性能显著优于基线方法
  • 支持未见过的任务与机器人形态,适合实际部署

强化学习代理在各类机器人任务中展现潜力,但仍严重依赖人工设计的奖励函数,需大量试错且常缺乏目标行为信息。本文提出REDS:基于分段示范的奖励学习框架,利用无动作视频与少量标注进行训练。通过将多源视频按子任务分割,并视其为真实奖励信号,训练一个以视频片段和对应子任务为条件的密集奖励函数,最小化等策略不变比较距离以确保对齐。同时采用对比学习对齐视频表示与子任务,提升在线交互中的子任务推断精度。实验表明,REDS在Meta-World的复杂操作任务及FurnitureBench的真实家具装配任务中均显著优于基线方法,且支持未见任务与机器人形态的泛化,具备大规模部署潜力。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) agents have demonstrated their potential across various robotic tasks. However, they still heavily rely on human-engineered reward functions, requiring extensive trial-and-error and access to target behavior information, often unavailable in real-world settings. This paper introduces REDS: REward learning from Demonstration with Segmentations, a novel reward learning framework that leverages action-free videos with minimal supervision. Specifically, REDS employs video demonstrations segmented into subtasks from diverse sources and treats these segments as ground-truth rewards. We train a dense reward function conditioned on video segments and their corresponding subtasks to ensure alignment with ground-truth reward signals by minimizing the Equivalent-Policy Invariant Comparison distance. Additionally, we employ contrastive learning objectives to align video representations with subtasks, ensuring precise subtask inference during online interactions. Our experiments show that REDS significantly outperforms baseline methods on complex robotic manipulation tasks in Meta-World and more challenging real-world tasks, such as furniture assembly in FurnitureBench, with minimal human intervention. Moreover, REDS facilitates generalization to unseen tasks and robot embodiments, highlighting its potential for scalable deployment in diverse environments.

奖励学习机器人子任务视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。