arXiv:2410.19115cs.CV2024-10CVPR被引 329

单目图像精准恢复3D几何结构,无需真实尺度标签。

MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

论文配图:MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
图 1 · 摘自论文原文
  • 采用仿射不变表示,消除训练中的尺度歧义。
  • 融合全局点云对齐与多尺度局部损失,提升几何精度。
  • 适用于开放域图像,适合需要高精度3D重建的场景。

我们提出MoGe,一种从单目开放域图像中恢复3D几何结构的强大模型。给定单张图像,该模型直接预测场景的3D点云图,采用仿射不变表示,对真实全局尺度和偏移无感。这一新表示避免了训练中的歧义监督,促进有效几何学习。此外,我们设计了一套新颖的全局与局部几何监督机制:包括一个鲁棒、最优且高效的点云对齐求解器,用于准确学习全局形状;以及多尺度局部几何损失,强化精细局部结构监督。模型在大规模混合数据集上训练,并在多样未见过的数据集上进行全面评估,显著优于当前最先进方法,在单目3D点云、深度图及相机视场估计任务中均表现优异。代码与模型可在项目页面获取。

原文摘要 · Abstract (English)

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation precludes ambiguous supervision in training and facilitate effective geometry learning. Furthermore, we propose a set of novel global and local geometry supervisions that empower the model to learn high-quality geometry. These include a robust, optimal, and efficient point cloud alignment solver for accurate global shape learning, and a multi-scale local geometry loss promoting precise local geometry supervision. We train our model on a large, mixed dataset and demonstrate its strong generalizability and high accuracy. In our comprehensive evaluation on diverse unseen datasets, our model significantly outperforms state-of-the-art methods across all tasks, including monocular estimation of 3D point map, depth map, and camera field of view. Code and models can be found on our project page.

3D重建单目估计几何学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。