arXiv:2512.19402cs.ROcs.CV2025-12中稿 · CVPR被引 8

用3D控制接口生成机器人操作新示范,大幅降低数据收集成本。

Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface

  • 通过3D点云编辑生成新操作轨迹,结合深度修正确保物理合理性。
  • 仅需1-5个原始示范,训练效果可媲美甚至超过50个真实示范。
  • 适合需要高效生成多样化机器人数据的研究者与开发者。

机器人学习的进展依赖大规模数据集和强大的视觉-运动策略架构,但策略鲁棒性仍受限于多样示范收集的巨大成本,尤其是在抓取任务中实现空间泛化。为减少重复数据采集,我们提出 Real2Edit2Real 框架,通过3D控制界面将3D可编辑性与2D视觉数据相连接。该方法首先利用多视角RGB观测重建场景几何,采用度量尺度的3D重建模型。基于重建几何,在点云上进行深度可靠的3D编辑,生成新的操作轨迹,并通过几何校正机器人姿态以恢复物理一致的深度,作为合成新示范的可靠条件。最后,我们提出一种以深度为主控信号的多条件视频生成模型,结合动作、边缘和射线图,合成空间增强的多视角操作视频。在四个真实世界操作任务上的实验表明,仅使用1-5个源示范生成的数据训练的策略,性能可达到甚至超越使用50个真实示范训练的策略,数据效率提升最高达10-50倍。此外,高度与纹理编辑实验验证了框架的灵活性与可扩展性,表明其有潜力成为统一的数据生成框架。

原文摘要 · Abstract (English)

Recent progress in robot learning has been driven by large-scale datasets and powerful visuomotor policy architectures, yet policy robustness remains limited by the substantial cost of collecting diverse demonstrations, particularly for spatial generalization in manipulation tasks. To reduce repetitive data collection, we present Real2Edit2Real, a framework that generates new demonstrations by bridging 3D editability with 2D visual data through a 3D control interface. Our approach first reconstructs scene geometry from multi-view RGB observations with a metric-scale 3D reconstruction model. Based on the reconstructed geometry, we perform depth-reliable 3D editing on point clouds to generate new manipulation trajectories while geometrically correcting the robot poses to recover physically consistent depth, which serves as a reliable condition for synthesizing new demonstrations. Finally, we propose a multi-conditional video generation model guided by depth as the primary control signal, together with action, edge, and ray maps, to synthesize spatially augmented multi-view manipulation videos. Experiments on four real-world manipulation tasks demonstrate that policies trained on data generated from only 1-5 source demonstrations can match or outperform those trained on 50 real-world demonstrations, improving data efficiency by up to 10-50x. Moreover, experimental results on height and texture editing demonstrate the framework's flexibility and extensibility, indicating its potential to serve as a unified data generation framework. Project website is https://real2edit2real.github.io/.

机器人学习数据生成3D编辑视觉-运动策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。