arXiv:2608.20788cs.CV2026-08

融合单目深度先验与多视角立体,提升深度图完整性和泛化能力。

M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo

论文配图:M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo
图 1 · 摘自论文原文
  • 双向互融:用多视角修正单目尺度,用单目补充多视角细节。
  • 在MVS benchmarks上超越现有方法,深度图更完整、边界更锐利。
  • 适合需要高精度与强泛化的3D重建场景,如稀疏视角下的应用。

基于深度学习的多视角立体(MVS)虽有显著进展,但在未见场景中泛化能力差,尤其在遮挡区域或视图重叠少的区域。现有方法将深度基础模型(DFM)作为单目深度先验引入MVS流程,但多采用静态单向融合,未能充分发挥双模态互补优势。本文提出新框架,通过双向互融策略将DFM与级联式MVS流水线紧密耦合:利用MVS深度解决单目预测中的尺度模糊问题,同时单目深度增强MVS估计的结构完整性和细粒度细节。此外,引入先验引导的成本体积优化机制,通过注意力融合和离散深度分桶有效整合多视角与单目信息,促进局部几何一致性。大量实验表明,本方法在标准基准上优于当前最佳MVS方法,生成更完整、更具泛化性的深度图,边界清晰。尽管未专为稀疏视角设计,仍表现出色,与专门的稀疏视角方法竞争激烈,且保持更优的精度-效率平衡。

原文摘要 · Abstract (English)

Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or regions with limited view overlap. To mitigate this, recent approaches integrate Depth Foundation Models (DFMs) into MVS pipelines to provide monocular depth priors. However, existing methods typically rely on a static, one-way fusion scheme, which fails to fully exploit the complementary strengths of both modalities. We propose a novel framework that overcomes this limitation by tightly coupling a DFM with a cascade MVS pipeline through a bidirectional mutual refinement strategy. Our method leverages MVS depth to resolve the scale ambiguity in monocular predictions, while the monocular depth, in turn, enhances the structural completeness and fine-grained detail of the MVS estimate. Furthermore, we introduce a prior-guided cost volume refinement mechanism that effectively integrates multi-view and monocular information via attention-based fusion and discretized depth bins, thereby promoting local geometric consistency. Extensive experiments demonstrate that our method outperforms state-of-the-art MVS approaches on standard benchmarks, producing more complete and generalizable depth maps with sharp boundaries. Furthermore, although not explicitly designed for sparse-view settings, our framework generalizes remarkably well, competing favorably with even dedicated sparse-view methods while maintaining a superior accuracy-efficiency trade-off.

三维重建深度估计多视角立体深度先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。