arXiv:2412.10778cs.CVcs.AI2024-12ICRA被引 5

无需奖励和标注,仅靠看视频就能学会复杂技能

Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos

  • 用自监督任务从无动作视频中推断专家行为
  • 在16个环境中实现领先性能,仅用视频即达成高精度策略
  • 适合缺乏标注数据或奖励信号的机器人学习场景

当前先进策略学习方法虽能生成专家级策略,但需依赖任务特定奖励、带动作标注的专家轨迹以及大量环境交互,成本高昂甚至不可行。相比之下,人类可通过观看少量互联网视频,在无额外监督下快速模仿学会技能。本文提出一种新框架UPESV,通过无监督方式从无动作视频中高效学习策略,无需奖励或任何其他专家监督。UPESV训练一个视频标注模型,利用多个有机融合的自监督任务,从专家视频中推断出专家动作。各任务协同工作,使模型充分利用无动作视频与无奖励交互,实现稳健的动力学理解与精准动作预测。同时,基于标注视频克隆策略,并收集环境交互用于自监督任务。经过样本高效的无监督迭代训练,最终获得基于鲁棒视频标注模型的先进策略。在16个挑战性程序生成环境中,UPESV在交互受限条件下表现优于5个先进基线(12/16任务领先),且仅依赖视频输入。

原文摘要 · Abstract (English)

Current advanced policy learning methodologies have demonstrated the ability to develop expert-level strategies when provided enough information. However, their requirements, including task-specific rewards, action-labeled expert trajectories, and huge environmental interactions, can be expensive or even unavailable in many scenarios. In contrast, humans can efficiently acquire skills within a few trials and errors by imitating easily accessible internet videos, in the absence of any other supervision. In this paper, we try to let machines replicate this efficient watching-and-learning process through Unsupervised Policy from Ensemble Self-supervised labeled Videos (UPESV), a novel framework to efficiently learn policies from action-free videos without rewards and any other expert supervision. UPESV trains a video labeling model to infer the expert actions in expert videos through several organically combined self-supervised tasks. Each task performs its duties, and they together enable the model to make full use of both action-free videos and reward-free interactions for robust dynamics understanding and advanced action prediction. Simultaneously, UPESV clones a policy from the labeled expert videos, in turn collecting environmental interactions for self-supervised tasks. After a sample-efficient, unsupervised, and iterative training process, UPESV obtains an advanced policy based on a robust video labeling model. Extensive experiments in sixteen challenging procedurally generated environments demonstrate that the proposed UPESV achieves state-of-the-art interaction-limited policy learning performance (outperforming five current advanced baselines on 12/16 tasks) without exposure to any other supervision except for videos.

无监督学习视频模仿策略克隆自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。