用3D一致性优化提升单目视频深度估计精度,解决帧间漂移问题。
3D Consistency Optimization for Self-Supervised Monocular Video Depth Estimation

- 将视频深度估计转为无约束多视角3D重建,利用3D先验统一建模。
- 在自监督和零样本临床场景中均达到领先精度,显著减少帧间漂移。
- 适合需要高精度3D感知的内窥镜导航等医疗视觉任务。
可靠的单目视频深度估计对内窥镜导航中的下游3D推理和具身AI至关重要。然而,现有自监督方法通常独立处理视频帧或依赖弱时序正则化,缺乏对潜在3D场景的整体感知,导致几何不一致预测和严重的跨帧漂移。为此,我们提出新范式,将序列视频深度估计重构为无约束多视角3D重建问题,充分挖掘近期3D基础模型中的强大几何先验。核心是基于三种约束的3D一致性优化框架:图像级光度渲染、显式世界坐标几何对齐以及多尺度时序梯度一致性。该统一优化巧妙地将孤立帧锚定于全局一致的3D结构中。方法在自监督训练场景及挑战性的零样本临床环境中均得到验证,结果表明所提方法在空间精度上优于基于帧的方法、基于视频的方法以及多视角3D重建基线。
原文摘要 · Abstract (English)
Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically treat video frames independently or rely on weak temporal regularization. These methods, lacking a holistic perception of the underlying 3D scene, inevitably suffer from geometrically inconsistent predictions and severe cross-frame drift. To address these limitations, we introduce a new paradigm that recasts sequential video depth estimation as an unconstrained multi-view 3D reconstruction problem, enabling full exploitation of the powerful geometric priors embedded in recent 3D foundation models. The core of our approach is a 3D consistency optimization framework driven by three constraints: image-level photometric rendering, explicit world-coordinate geometric alignment, and multi-scale temporal gradient consistency. Such unified optimization elegantly anchors isolated frames to a globally coherent 3D structure. Our method has been validated in both the self-supervised training scenarios and challenging zero-shot clinical environments. Results show that the proposed approach achieves state-of-the-art spatial accuracy, outperforming the frame-based, video-based depth estimators and the multi-view 3D reconstruction baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。