arXiv:2503.22430cs.CV2025-03CVPR被引 35

无需训练即可跨场景精准估算深度,解决多视角立体匹配泛化难题。

MVSAnywhere: Zero-Shot Multi-View Stereo

  • 融合单目与多视角线索,自适应构建成本体积应对尺度差异。
  • 在鲁棒多视角深度基准上达到顶尖零样本性能,优于现有方法。
  • 适合需要快速部署、跨域通用的3D重建与自动驾驶应用。

从多视角精确计算深度是计算机视觉中的基础性长期挑战。然而,现有大多数方法在不同领域和场景类型(如室内与室外)之间泛化能力差。训练通用多视角立体模型面临诸多问题,例如如何最优利用基于Transformer的架构、如何处理输入视角数量不固定时的附加元数据,以及如何估计有效深度范围——该范围在不同场景间差异显著且通常事先未知。为此,我们提出MVSA,一种新颖且通用的多视角立体架构,旨在实现“随处可用”:跨多种领域和深度范围泛化。MVSA结合单目与多视角线索,并采用自适应成本体积以应对尺度相关问题。我们在鲁棒多视角深度基准上展示了最先进的零样本深度估计性能,超越了现有的多视角立体与单目基线方法。

原文摘要 · Abstract (English)

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision. However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs. outdoor). Training a general-purpose multi-view stereo model is challenging and raises several questions, e.g. how to best make use of transformer-based architectures, how to incorporate additional metadata when there is a variable number of input views, and how to estimate the range of valid depths which can vary considerably across different scenes and is typically not known a priori? To address these issues, we introduce MVSA, a novel and versatile Multi-View Stereo architecture that aims to work Anywhere by generalizing across diverse domains and depth ranges. MVSA combines monocular and multi-view cues with an adaptive cost volume to deal with scale-related issues. We demonstrate state-of-the-art zero-shot depth estimation on the Robust Multi-View Depth Benchmark, surpassing existing multi-view stereo and monocular baselines.

多视角立体零样本自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。