用3D编辑+2D视频补全,让少样本抓取数据更通用。
R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies

- 先在3D空间改物体和机械臂轨迹,再投影到2D视频补全画面。
- 在少样本下提升2D策略对空间变化的适应能力,效果优于基线。
- 适合做少样本机器人抓取、视觉-语言-动作联合学习的研究者。
空间泛化对模仿学习的抓取策略至关重要,但通常需在多种物体姿态、机器人配置和相机视角下扩展演示数据,成本高昂。从少量源演示中进行数据增强是更具性价比的替代方案。基于仿真的增强可控制变化,但需复杂环境与物体设置,且存在仿真到现实的差距。近期真实世界到真实世界的增强方法通过联合编辑真实演示中的3D观测与动作轨迹来规避这些问题,但仍依赖强3D场景解析与几何补全,且常生成适配3D点云策略而非基于RGB的2D策略的观测。我们提出R2RDreamer,一种真实世界到真实世界的演示增强框架,在保持3D动作-观测编辑几何一致性的同时,将视觉补全过程移至2D视频空间。具体而言,R2RDreamer首先在共享3D坐标系中轻量级地编辑不完整物体点云与末端执行器轨迹;随后将编辑后的场景投影至带掩码的图像空间控制视频,并通过遮挡感知推理,利用密集控制的图像到视频模型完成时序一致的RGB观测补全。在多个空间偏移的抓取任务上,使用2D扩散型策略与视觉-语言-动作策略的实验表明,R2RDreamer能显著提升有限源演示下的空间泛化能力,分析验证了3D编辑、遮挡感知投影与视频补全的贡献。
原文摘要 · Abstract (English)
Spatial generalization is critical for imitation-learned manipulation policies, but achieving it typically requires scaling demonstrations across diverse object poses, robot configurations, and camera viewpoints. Data augmentation from a few source demonstrations offers a practical alternative to costly real-world collection. Simulation-based augmentation can create controllable variation, but requires complex environment and object setup and may introduce a sim-to-real gap. Recent real-to-real methods avoid these issues by jointly editing 3D observations and action trajectories from real demonstrations, yet they still rely on strong 3D scene parsing and geometry completion, and often produce observations tailored to 3D pointcloud policies rather than RGB-based 2D policies. We propose R2RDreamer, a real-to-real demonstration augmentation framework that preserves the geometric consistency of 3D action-observation editing while moving visual completion to 2D video space. Specifically, R2RDreamer first performs lightweight 3D augmentation by editing incomplete object pointclouds and end-effector trajectories in a shared 3D frame; it then projects the edited scene into masked image-space control videos with occlusion-aware reasoning and uses a dense-control image-to-video model to complete temporally coherent RGB observations. Experiments on spatially shifted manipulation tasks with both 2D diffusion-style policies and vision-language-action policies show that R2RDreamer improves spatial generalization from limited source demonstrations, with analyses validating the contributions of 3D editing, occlusion-aware projection, and video completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。