通过融合视内与视间关系,提升多视角立体匹配精度
ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
- 设计视内特征融合模块,利用单图特征坐标相关性增强匹配
- 引入轻量级视间聚合模块,通过体素上下文关联引导正则化
- 在DTU和Tanks&Temples上表现优于主流方法,且计算开销更低
多视角立体(MVS)旨在从一系列重叠图像中估计深度并重建三维点云。现有基于学习的MVS框架忽视了特征中嵌入的几何信息与相关性,导致代价匹配能力弱。本文提出ICG-MVSNet,显式整合视内与视间关系以指导深度估计。具体而言,我们设计了一个视内特征融合模块,利用单张图像内特征坐标的关联性来增强鲁棒的代价匹配;同时引入一个轻量级的视间聚合模块,高效利用体素间的上下文信息来引导正则化。该方法在DTU数据集和Tanks and Temples基准上均取得与当前最优方法相当的性能,且所需计算资源更少。
原文摘要 · Abstract (English)
Multi-view Stereo (MVS) aims to estimate depth and reconstruct 3D point clouds from a series of overlapping images. Recent learning-based MVS frameworks overlook the geometric information embedded in features and correlations, leading to weak cost matching. In this paper, we propose ICG-MVSNet, which explicitly integrates intra-view and cross-view relationships for depth estimation. Specifically, we develop an intra-view feature fusion module that leverages the feature coordinate correlations within a single image to enhance robust cost matching. Additionally, we introduce a lightweight cross-view aggregation module that efficiently utilizes the contextual information from volume correlations to guide regularization. Our method is evaluated on the DTU dataset and Tanks and Temples benchmark, consistently achieving competitive performance against state-of-the-art works, while requiring lower computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。