提出无需校准的3D重建模型,适配广角镜头成像。
CAM3R: Camera-Agnostic Model for 3D Reconstruction
- 分路设计射线模块与视图交互模块,联合估计方向与距离。
- 在全景、鱼眼等多类镜头数据上实现最优姿态与重建精度。
- 无需相机标定,适用于非标准光学设备,适合实际部署。
从无姿态图像中恢复稠密三维几何仍是计算机视觉的基础挑战。现有最先进模型主要基于透视数据集训练,隐式依赖标准针孔相机模型,因此在鱼眼或全景传感器采集的广角影像上会出现显著几何退化。为此,我们提出CAM3R——一种无需相机标定即可处理广角相机模型的前馈式3D重建模型。其框架包含双视图网络,分为射线模块(RM)用于估计每像素射线方向,以及跨视图模块(CVM)用于推断径向距离与置信度图、点图和相对位姿。为将这些局部预测统一为一致的3D场景,我们引入射线感知全局对齐框架,实现位姿精修与尺度优化,同时严格保留局部几何结构。在包括全景、鱼眼和针孔在内的多种相机模型数据集上的大量实验表明,CAM3R在姿态估计与重建任务上均达到新基准性能。
原文摘要 · Abstract (English)
Recovering dense 3D geometry from unposed images remains a foundational challenge in computer vision. Current state-of-the-art models are predominantly trained on perspective datasets, which implicitly constrains them to a standard pinhole camera geometry. As a result, these models suffer from significant geometric degradation when applied to wide-angle imagery captured via non-rectilinear optics, such as fisheye or panoramic sensors. To address this, we present CAM3R, a Camera-Agnostic, feed-forward Model for 3D Reconstruction capable of processing images from wide-angle camera models without prior calibration. Our framework consists of a two-view network which is bifurcated into a Ray Module (RM) to estimate per-pixel ray directions and a Cross-view Module (CVM) to infer radial distance with confidence maps, pointmaps, and relative poses. To unify these pairwise predictions into a consistent 3D scene, we introduce a Ray-Aware Global Alignment framework for pose refinement and scale optimization while strictly preserving the predicted local geometry. Extensive experiments on various camera model datasets, including panorama, fisheye and pinhole imagery, demonstrate that CAM3R establishes a new state-of-the-art in pose estimation and reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。