从2D视角学习通用3D场景表示,让机器人直接理解全局空间。
Towards Learning a Generalizable 3D Scene Representation from 2D Observations
- 在全局坐标系中构建3D占用表示,无需依赖相机视角
- 40个真实场景训练,重建误差仅26mm,含遮挡区域
- 无需微调即可适应新物体布局,适合机器人操作任务
我们提出一种可泛化的神经辐射场方法,从机器人第一人称视角观测中预测3D工作空间占用情况。与以往基于相机中心坐标的模型不同,该模型在全局工作空间帧中构建占用表示,可直接用于机器人操作。模型融合灵活的源视角,无需针对特定场景微调即可泛化到未见过的物体排列。我们在人形机器人上验证了该方法,并将预测几何与3D传感器真值对比。模型在40个真实场景上训练,实现26mm重建误差,包含遮挡区域,验证了其超越传统立体视觉方法推断完整3D占用的能力。
原文摘要 · Abstract (English)
We introduce a Generalizable Neural Radiance Field approach for predicting 3D workspace occupancy from egocentric robot observations. Unlike prior methods operating in camera-centric coordinates, our model constructs occupancy representations in a global workspace frame, making it directly applicable to robotic manipulation. The model integrates flexible source views and generalizes to unseen object arrangements without scene-specific finetuning. We demonstrate the approach on a humanoid robot and evaluate predicted geometry against 3D sensor ground truth. Trained on 40 real scenes, our model achieves 26mm reconstruction error, including occluded regions, validating its ability to infer complete 3D occupancy beyond traditional stereo vision methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。