单图恢复度量3D形状,同时估计深度与相机内参。
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery
- 深度与相机内参互为约束,联合优化提升精度。
- 在多个室内室外数据集上表现优异,3D形状质量显著提升。
- 适合机器人导航与交互场景中的空间感知需求。
从单张图像恢复度量3D形状对机器人和具身智能应用至关重要,但仅靠深度无法恢复真实尺度,需相机内参配合。本文理论证明深度可作为相机内参估计的3D先验,并揭示两者间的相互依赖关系。为此提出协同学习框架CoL3D,统一网络在深度、相机内参和3D点云三个层面进行联合优化。设计规范入射场机制作为先验,增强内参校准;引入点云空间的形状相似性损失,提升3D形状质量。在多个室内与室外基准数据集(in-domain)上训练与测试,CoL3D在深度估计与相机标定任务中均表现突出,显著提升机器人感知能力所需的3D形状精度。
原文摘要 · Abstract (English)
Recovering the metric 3D shape from a single image is particularly relevant for robotics and embodied intelligence applications, where accurate spatial understanding is crucial for navigation and interaction with environments. Usually, the mainstream approaches achieve it through monocular depth estimation. However, without camera intrinsics, the 3D metric shape can not be recovered from depth alone. In this study, we theoretically demonstrate that depth serves as a 3D prior constraint for estimating camera intrinsics and uncover the reciprocal relations between these two elements. Motivated by this, we propose a collaborative learning framework for jointly estimating depth and camera intrinsics, named CoL3D, to learn metric 3D shapes from single images. Specifically, CoL3D adopts a unified network and performs collaborative optimization at three levels: depth, camera intrinsics, and 3D point clouds. For camera intrinsics, we design a canonical incidence field mechanism as a prior that enables the model to learn the residual incident field for enhanced calibration. Additionally, we incorporate a shape similarity measurement loss in the point cloud space, which improves the quality of 3D shapes essential for robotic applications. As a result, when training and testing on a single dataset with in-domain settings, CoL3D delivers outstanding performance in both depth estimation and camera calibration across several indoor and outdoor benchmark datasets, which leads to remarkable 3D shape quality for the perception capabilities of robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。