arXiv:2506.05655cs.CVeess.IV2025-06被引 2

针对航拍三维重建难题,提出自适应深度范围与法向量引导的新方法。

Aerial Multi-View Stereo via Adaptive Depth Range Inference and Normal Cues

  • 通过交叉注意力差异学习生成自适应深度范围图,突破固定深度边界限制。
  • 在WHU、LuoJia-MVS和München数据集上达到当前最优精度,计算开销更低。
  • 适合需要高精度航拍三维建模的地理信息、智慧城市等应用场景。

从多视角航拍图像进行三维数字城市重建是一项关键应用,深度多视图立体(MVS)方法已超越传统技术。然而,现有方法常忽视航拍与近距离场景的关键差异,如沿视差线变化的深度范围及低细节航拍图像导致的特征匹配不敏感问题。为此,我们提出自适应深度范围MVS(ADR-MVS),融合单目几何线索以提升多视图深度估计精度。核心是深度范围预测器,通过交叉注意力差异学习,利用深度和法向量估计生成自适应范围图。第一阶段,基于单目线索的范围图打破预设深度边界,增强特征匹配判别力并避免陷入局部最优;后续阶段,推断范围逐步收缩,最终适配级联MVS框架实现精确深度回归。此外,设计了法向量引导的成本聚合操作,提升代价体中的几何感知能力。最后引入法向量引导的深度精修模块,优于现有RGB引导技术。实验表明,ADR-MVS在WHU、LuoJia-MVS和München数据集上均达当前最优性能,且具有更优的计算复杂度。

原文摘要 · Abstract (English)

Three-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly overlook the key differences between aerial and close-range settings, such as varying depth ranges along epipolar lines and insensitive feature-matching associated with low-detailed aerial images. To address these issues, we propose an Adaptive Depth Range MVS (ADR-MVS), which integrates monocular geometric cues to improve multi-view depth estimation accuracy. The key component of ADR-MVS is the depth range predictor, which generates adaptive range maps from depth and normal estimates using cross-attention discrepancy learning. In the first stage, the range map derived from monocular cues breaks through predefined depth boundaries, improving feature-matching discriminability and mitigating convergence to local optima. In later stages, the inferred range maps are progressively narrowed, ultimately aligning with the cascaded MVS framework for precise depth regression. Moreover, a normal-guided cost aggregation operation is specially devised for aerial stereo images to improve geometric awareness within the cost volume. Finally, we introduce a normal-guided depth refinement module that surpasses existing RGB-guided techniques. Experimental results demonstrate that ADR-MVS achieves state-of-the-art performance on the WHU, LuoJia-MVS, and München datasets, while exhibits superior computational complexity.

三维重建航拍图像深度估计多视图立体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。