让视频模型通过自探索学会执行动作,无需人工标注。
Grounding Video Models to Actions through Goal Conditioned Exploration
- 用生成的视频画面作目标,让智能体自主探索学习动作
- 在8个Libero、6个MetaWorld等任务上达到或超越专家示范效果
- 无需奖励、动作标签或分割掩码,适合零样本迁移场景
大规模视频模型虽蕴含丰富的物理动态知识,但缺乏与智能体行为的关联,无法指导如何实现视频中的视觉状态。现有方法依赖特定数据训练视觉逆动力学模型,成本高且泛化受限。本文提出一种框架,通过在具身环境中进行自探索,利用生成的视频状态作为视觉目标,结合轨迹级动作生成与视频引导,使智能体在无外部监督(如奖励、动作标签、分割掩码)条件下完成复杂任务。我们在Libero(8项)、MetaWorld(6项)、Calvin(4项)和iThor视觉导航(12项)上验证该方法,结果表明其性能媲美甚至超过多个基于专家演示的行为克隆基线。
原文摘要 · Abstract (English)
Large video models, pretrained on massive amounts of Internet video, provide a rich source of physical knowledge about the dynamics and motions of objects and tasks. However, video models are not grounded in the embodiment of an agent, and do not describe how to actuate the world to reach the visual states depicted in a video. To tackle this problem, current methods use a separate vision-based inverse dynamic model trained on embodiment-specific data to map image states to actions. Gathering data to train such a model is often expensive and challenging, and this model is limited to visual settings similar to the ones in which data are available. In this paper, we investigate how to directly ground video models to continuous actions through self-exploration in the embodied environment -- using generated video states as visual goals for exploration. We propose a framework that uses trajectory level action generation in combination with video guidance to enable an agent to solve complex tasks without any external supervision, e.g., rewards, action labels, or segmentation masks. We validate the proposed approach on 8 tasks in Libero, 6 tasks in MetaWorld, 4 tasks in Calvin, and 12 tasks in iThor Visual Navigation. We show how our approach is on par with or even surpasses multiple behavior cloning baselines trained on expert demonstrations while without requiring any action annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。