用开放世界图片生成机器人可执行的视觉动作数据,实现低成本大规模训练
IGen: Scalable Data Generation for Robot Learning from Open-World Images
- 将2D图片转为3D场景,用视觉语言模型生成任务计划和机械臂动作序列
- 合成动态场景演化与连贯视觉观测,生成的数据让机器人策略性能接近真实数据
- 适合需要海量训练数据的通用机器人学习研究者使用
通用机器人策略的兴起带来了对大规模训练数据的指数级需求。然而,机器人实机数据采集耗时费力,且通常局限于特定环境。相比之下,开放世界图像包含丰富多样的真实场景,天然契合机器人操作任务,为低成本、大规模机器人数据获取提供了可能。但缺乏关联的机器人动作,使这些图像难以直接用于机器人学习,导致这一丰富视觉资源长期未被利用。为此,我们提出IGen框架,可从开放世界图像中可扩展地生成逼真的视觉观测与可执行动作。IGen首先将无结构的2D像素转化为适用于场景理解与操作的结构化3D表示;随后利用视觉语言模型将场景特定任务指令转化为高层规划,并生成低层动作(作为SE(3)末端位姿序列)。基于这些位姿,合成动态场景演化并渲染时间连贯的视觉观测。实验验证了IGen生成的视觉-运动数据质量高,仅用IGen合成数据训练的策略性能可媲美真实数据训练的策略。这表明IGen有望支持从开放世界图像中大规模生成数据,助力通用机器人策略训练。
原文摘要 · Abstract (English)
The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-intensive and often limited to specific environments. In contrast, open-world images capture a vast diversity of real-world scenes that naturally align with robotic manipulation tasks, offering a promising avenue for low-cost, large-scale robot data acquisition. Despite this potential, the lack of associated robot actions hinders the practical use of open-world images for robot learning, leaving this rich visual resource largely unexploited. To bridge this gap, we propose IGen, a framework that scalably generates realistic visual observations and executable actions from open-world images. IGen first converts unstructured 2D pixels into structured 3D scene representations suitable for scene understanding and manipulation. It then leverages the reasoning capabilities of vision-language models to transform scene-specific task instructions into high-level plans and generate low-level actions as SE(3) end-effector pose sequences. From these poses, it synthesizes dynamic scene evolution and renders temporally coherent visual observations. Experiments validate the high quality of visuomotor data generated by IGen, and show that policies trained solely on IGen-synthesized data achieve performance comparable to those trained on real-world data. This highlights the potential of IGen to support scalable data generation from open-world images for generalist robotic policy training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。