只需一张照片,就能生成可交互的机器人训练环境。
Robot Learning from Any Images
- 从单张图像恢复物理场景,直接生成机器人可用数据。
- 几分钟内从各类图片生成海量视觉-动作示范数据。
- 适合想快速构建机器人训练数据的研究者和开发者。
我们提出RoLA框架,将任意真实场景图像转化为可交互、具备物理特性的机器人环境。与以往方法不同,RoLA仅需单张图像,无需额外硬件或数字资产。该框架通过单视角物理场景重建与高效视觉融合策略,实现从相机拍摄、机器人数据集及网络图片等多源图像中,快速生成大量视觉-动作示范数据。整个过程可在几分钟内完成,显著降低数据生成门槛。我们在多个应用场景验证了其通用性,包括可扩展的机器人数据生成与增强、基于互联网图像的机器人学习,以及针对机械臂和人形机器人的单图真实-仿真-真实系统。视频演示见https://sihengz02.github.io/RoLA。
原文摘要 · Abstract (English)
We introduce RoLA, a framework that transforms any in-the-wild image into an interactive, physics-enabled robotic environment. Unlike previous methods, RoLA operates directly on a single image without requiring additional hardware or digital assets. Our framework democratizes robotic data generation by producing massive visuomotor robotic demonstrations within minutes from a wide range of image sources, including camera captures, robotic datasets, and Internet images. At its core, our approach combines a novel method for single-view physical scene recovery with an efficient visual blending strategy for photorealistic data collection. We demonstrate RoLA's versatility across applications like scalable robotic data generation and augmentation, robot learning from Internet images, and single-image real-to-sim-to-real systems for manipulators and humanoids. Video results are available at https://sihengz02.github.io/RoLA .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。