arXiv:2501.08643cs.CV2025-01TPAMI被引 34

融合单目深度先验,统一解决立体匹配与多视角立体重建难题。

MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors

  • 构建双分支架构,融合单目与多视角深度信息,迭代优化几何估计。
  • 在8个基准上达到新最好性能,实时版本显著超越现有实时方法。
  • 适合需要高精度、强泛化能力的三维重建与自动驾驶场景使用。

我们提出MonSter++,一种用于多视角深度估计的几何基础模型,统一了校正后立体匹配与未校正多视角立体重建。两者均通过对应关系搜索恢复度量深度,面临相同困境:在匹配线索不足的区域难以处理。为此,我们提出将单目深度先验融入多视角深度估计,有效结合单视角与多视角信息的优势。MonSter++采用双分支结构融合单目与多视角深度,基于置信度引导自适应选择可靠多视角线索,纠正单目深度的尺度模糊性;而优化后的单目预测则反过来指导多视角估计在困难区域的表现。这种迭代相互增强机制使MonSter++能从粗粒度对象级单目先验演化为细粒度像素级几何结构,充分释放多视角深度估计潜力。MonSter++在立体匹配与多视角立体重建任务上均取得新最优结果。通过级联搜索与多尺度深度融合策略,其实时版本RT-MonSter++也大幅领先于此前实时方法。如图1所示,其在三个任务共八个基准上显著优于以往方法,展现出强大通用性。此外,该模型还具备优异的零样本泛化能力。我们将公开大模型与实时模型,以支持开源社区使用。

原文摘要 · Abstract (English)

We introduce MonSter++, a geometric foundation model for multi-view depth estimation, unifying rectified stereo matching and unrectified multi-view stereo. Both tasks fundamentally recover metric depth from correspondence search and consequently face the same dilemma: struggling to handle ill-posed regions with limited matching cues. To address this, we propose MonSter++, a novel method that integrates monocular depth priors into multi-view depth estimation, effectively combining the complementary strengths of single-view and multi-view cues. MonSter++ fuses monocular depth and multi-view depth into a dual-branched architecture. Confidence-based guidance adaptively selects reliable multi-view cues to correct scale ambiguity in monocular depth. The refined monocular predictions, in turn, effectively guide multi-view estimation in ill-posed regions. This iterative mutual enhancement enables MonSter++ to evolve coarse object-level monocular priors into fine-grained, pixel-level geometry, fully unlocking the potential of multi-view depth estimation. MonSter++ achieves new state-of-the-art on both stereo matching and multi-view stereo. By effectively incorporating monocular priors through our cascaded search and multi-scale depth fusion strategy, our real-time variant RT-MonSter++ also outperforms previous real-time methods by a large margin. As shown in Fig.1, MonSter++ achieves significant improvements over previous methods across eight benchmarks from three tasks -- stereo matching, real-time stereo matching, and multi-view stereo, demonstrating the strong generality of our framework. Besides high accuracy, MonSter++ also demonstrates superior zero-shot generalization capability. We will release both the large and the real-time models to facilitate their use by the open-source community.

深度估计立体匹配多视角重建单目先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。