用视觉身份提示生成多视角操控视频,提升机器人训练数据质量
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
- 用示例图像作为视觉提示,精准控制场景布局生成
- 在仿真和真实机器人上均实现策略性能显著提升
- 适合需要高质量多视角操控数据的研究者
操控数据的多样性、数量和质量对训练有效机器人策略至关重要。但受硬件与物理环境限制,大规模真实世界操控数据难以扩展至多样场景。近期工作利用文本提示引导的图像扩散模型,通过改变视觉观测中的背景和桌面上物体来扩充数据。然而,这些方法常忽视先进策略模型所需的多视角和时序一致性观测。此外,仅靠文本提示无法可靠指定场景配置。为此,我们提出视觉身份提示,将示例图像作为条件输入,引导生成期望的场景布局。同时构建可扩展的流水线,从大规模机器人数据集中构建视觉身份库。使用该增强数据训练下游视觉-语言-动作及视觉运动策略模型,在仿真与真实机器人环境中均取得稳定性能提升。
原文摘要 · Abstract (English)
The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to scale across diverse environments. Recent work uses text-prompt conditioned image diffusion models to augment manipulation data by altering the backgrounds and tabletop objects in the visual observations. However, these approaches often overlook the practical need for multi-view and temporally coherent observations required by state-of-the-art policy models. Further, text prompts alone cannot reliably specify the scene setup. To provide the diffusion model with explicit visual guidance, we introduce visual identity prompting, which supplies exemplar images as conditioning inputs to guide the generation of the desired scene setup. To this end, we also build a scalable pipeline to curate a visual identity pool from large robotics datasets. Using our augmented manipulation data to train downstream vision-language-action and visuomotor policy models yields consistent performance gains in both simulation and real-robot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。