用视频生成模型自动生成机器人任务训练视频,提升复杂任务学习效果。
LuciBot: Automated Robot Policy Learning from Generated Videos
- 用视频生成模型根据文本描述生成任务完成视频。
- 从视频中提取6D姿态、2D分割和深度等丰富监督信号。
- 适合需要复杂视觉理解的机器人仿真训练场景。
自动生成具身任务的训练监督至关重要,因为手动设计既繁琐又不可扩展。以往方法使用大语言模型(LLMs)或视觉-语言模型(VLMs)生成奖励,但主要局限于奖励定义清晰的简单任务(如抓取放置)。这是因为LLMs难以解析压缩为文本或代码的复杂场景,而基于VLM的奖励受限于输出表达能力。为此,我们利用通用视频生成模型的想象力:给定初始仿真帧和文本任务描述,模型生成语义正确的任务完成视频。我们从中提取丰富的监督信号,包括6D物体姿态序列、2D分割图和估计深度,用于仿真环境中的任务学习。该方法显著提升了复杂具身任务的监督质量,支持大规模模拟训练。
原文摘要 · Abstract (English)
Automatically generating training supervision for embodied tasks is crucial, as manual designing is tedious and not scalable. While prior works use large language models (LLMs) or vision-language models (VLMs) to generate rewards, these approaches are largely limited to simple tasks with well-defined rewards, such as pick-and-place. This limitation arises because LLMs struggle to interpret complex scenes compressed into text or code due to their restricted input modality, while VLM-based rewards, though better at visual perception, remain limited by their less expressive output modality. To address these challenges, we leverage the imagination capability of general-purpose video generation models. Given an initial simulation frame and a textual task description, the video generation model produces a video demonstrating task completion with correct semantics. We then extract rich supervisory signals from the generated video, including 6D object pose sequences, 2D segmentations, and estimated depth, to facilitate task learning in simulation. Our approach significantly improves supervision quality for complex embodied tasks, enabling large-scale training in simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。