用RGB图像和机器人状态实时生成高精度3D场景,提升机械臂操作能力。
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
- 从RGB图和机器人状态直接推断三维几何,无需依赖深度传感器。
- 在400万帧合成数据上训练,重建精度超越现有方法和真实深度传感器。
- 适合需要精准3D感知的机器人抓取、仿真到现实迁移等任务。
3D空间感知是通用机器人操作的基础,但获取可靠且高质量的3D几何仍具挑战。深度传感器易受噪声和材质影响,而现有重建模型缺乏物理交互所需的精度与度量一致性。我们提出Robo3R,一种前馈式、适用于操作的3D重建模型,可实时从RGB图像和机器人状态中直接预测精确的度量尺度场景几何。Robo3R联合推断尺度不变的局部几何与相对相机位姿,并通过学习的全局相似变换统一至标准机器人坐标系。为满足操作精度需求,采用掩码点头生成清晰细粒度点云,并使用基于关键点的透视-三点(PnP)方法优化相机外参与全局对齐。在包含四百万帧高保真标注数据的Robo3R-4M合成数据集上训练,Robo3R持续优于当前最优重建方法及深度传感器。在模仿学习、仿真到现实迁移、抓取生成和无碰撞运动规划等下游任务中均取得一致性能提升,表明该3D感知模块在机器人操作中的巨大潜力。
原文摘要 · Abstract (English)
3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for physical interaction. We introduce Robo3R, a feed-forward, manipulation-ready 3D reconstruction model that predicts accurate, metric-scale scene geometry directly from RGB images and robot states in real time. Robo3R jointly infers scale-invariant local geometry and relative camera poses, which are unified into the scene representation in the canonical robot frame via a learned global similarity transformation. To meet the precision demands of manipulation, Robo3R employs a masked point head for sharp, fine-grained point clouds, and a keypoint-based Perspective-n-Point (PnP) formulation to refine camera extrinsics and global alignment. Trained on Robo3R-4M, a curated large-scale synthetic dataset with four million high-fidelity annotated frames, Robo3R consistently outperforms state-of-the-art reconstruction methods and depth sensors. Across downstream tasks including imitation learning, sim-to-real transfer, grasp synthesis, and collision-free motion planning, we observe consistent gains in performance, suggesting the promise of this alternative 3D sensing module for robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。