arXiv:2504.11845cs.CV2025-04

用深度基础模型生成先验,让无真实标签的多视图立体重建更准。

Boosting Multi-View Stereo with Depth Foundation Model in the Absence of Real-World Labels

  • 用深度基础模型生成深度先验,模拟真实立体对应关系
  • 在DTU和Tanks & Temples上超越现有无标签方法
  • 适合无真实标注数据的3D重建场景

近年来基于学习的多视图立体(MVS)方法取得了显著进展。然而,如何在不使用真实世界标签的情况下有效训练网络仍是难题。本文受视觉基础模型进展启发,提出一种新方法DFM-MVS,利用深度基础模型生成有效的深度先验,以提升无真实标签条件下的MVS性能。具体地,设计了一种基于深度先验的伪监督训练机制,通过生成的深度先验模拟真实的立体对应关系,从而为MVS网络构建有效监督;此外,提出一种深度先验引导的误差修正策略,利用深度先验指导,缓解广泛采用的粗到精网络结构中的误差传播问题。在DTU和Tanks & Temples数据集上的实验结果表明,所提DFM-MVS显著优于现有无真实标签的MVS方法。

原文摘要 · Abstract (English)

Learning-based Multi-View Stereo (MVS) methods have made remarkable progress in recent years. However, how to effectively train the network without using real-world labels remains a challenging problem. In this paper, driven by the recent advancements of vision foundation models, a novel method termed DFM-MVS, is proposed to leverage the depth foundation model to generate the effective depth prior, so as to boost MVS in the absence of real-world labels. Specifically, a depth prior-based pseudo-supervised training mechanism is developed to simulate realistic stereo correspondences using the generated depth prior, thereby constructing effective supervision for the MVS network. Besides, a depth prior-guided error correction strategy is presented to leverage the depth prior as guidance to mitigate the error propagation problem inherent in the widely-used coarse-to-fine network structure. Experimental results on DTU and Tanks & Temples datasets demonstrate that the proposed DFM-MVS significantly outperforms existing MVS methods without using real-world labels.

多视图立体深度先验无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。