单目视频实现大场景毫米级精准重建,突破深度模糊、位姿漂移难题。
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
- 联合优化深度、位姿与局部辐射场,避免孤立求解的误差累积。
- 在八组场景上绝对轨迹误差低至0.001-0.021米,较基准方法降低18倍。
- 支持城市街区级覆盖,仅用一张显卡即可完成增量式建模。
从单目视频进行逼真三维重建在大规模场景中会因深度、位姿和辐射场孤立求解而失效:尺度模糊导致鬼影几何,长时程位姿漂移破坏对齐,单一全局NeRF无法建模数百米内容。本文提出一种联合学习框架,耦合三者并有效克服各失败案例。系统首先使用带度量尺度监督的视觉变换器(ViT)深度网络,实现视场变化下的全局一致深度。多尺度特征束调整(BA)层在特征空间直接优化相机位姿,利用学习到的金字塔描述符替代脆弱的关键点,抑制无约束轨迹中的漂移。对于场景表示,采用增量式局部辐射场层级结构:当视角重叠低于阈值时,即时分配并冻结新的哈希网格NeRF,实现在单张显卡上对街区级范围的覆盖。在Tanks and Temples基准测试中,本方法在八组室内外序列上将绝对轨迹误差降至0.001–0.021米,相比BARF降低达18倍,比NoPe-NeRF降低2倍,同时保持亚像素级相对位姿误差。结果表明,仅用一台未标定的RGB相机即可实现米级精度的三维重建与高保真新视角合成。
原文摘要 · Abstract (English)
Photorealistic 3-D reconstruction from monocular video collapses in large-scale scenes when depth, pose, and radiance are solved in isolation: scale-ambiguous depth yields ghost geometry, long-horizon pose drift corrupts alignment, and a single global NeRF cannot model hundreds of metres of content. We introduce a joint learning framework that couples all three factors and demonstrably overcomes each failure case. Our system begins with a Vision-Transformer (ViT) depth network trained with metric-scale supervision, giving globally consistent depths despite wide field-of-view variations. A multi-scale feature bundle-adjustment (BA) layer refines camera poses directly in feature space--leveraging learned pyramidal descriptors instead of brittle keypoints--to suppress drift on unconstrained trajectories. For scene representation, we deploy an incremental local-radiance-field hierarchy: new hash-grid NeRFs are allocated and frozen on-the-fly when view overlap falls below a threshold, enabling city-block-scale coverage on a single GPU. Evaluated on the Tanks and Temples benchmark, our method reduces Absolute Trajectory Error to 0.001-0.021 m across eight indoor-outdoor sequences--up to 18x lower than BARF and 2x lower than NoPe-NeRF--while maintaining sub-pixel Relative Pose Error. These results demonstrate that metric-scale, drift-free 3-D reconstruction and high-fidelity novel-view synthesis are achievable from a single uncalibrated RGB camera.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。