arXiv:2603.27542cs.CV2026-03中稿 · ed被引 3

多视角匹配新方法,提升3D重建的稠密与准确性

MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction

  • 利用成对匹配结果作为几何先验,构建高效多视角匹配架构
  • 在多个基准上实现更稠密、更准确的3D重建,优于现有稀疏与稠密匹配方法
  • 适合需要高精度3D重建的视觉任务,如运动恢复结构(SfM)

在三维视觉任务中,如运动恢复结构(SfM),建立图像间的稳定对应关系至关重要。然而,现有匹配器多为成对操作,在跨视图链式连接时常产生碎片化且几何不一致的轨迹。本文提出MV-RoMa,一种多视角稠密匹配模型,可联合估计源图像到多个共可见目标的稠密对应关系。具体而言,设计了高效模型架构:(i) 多视角编码器,利用成对匹配结果作为几何先验;(ii) 多视角匹配精修模块,通过像素级注意力优化对应关系。此外,提出后处理策略,将模型生成的一致多视角对应关系整合为高质量轨迹用于SfM。在多样且具有挑战性的基准测试中,MV-RoMa生成的对应关系更可靠,3D重建更稠密、更准确,显著优于现有稀疏与稠密匹配方法。

原文摘要 · Abstract (English)

Establishing consistent correspondences across images is essential for 3D vision tasks such as structure-from-motion (SfM), yet most existing matchers operate in a pairwise manner, often producing fragmented and geometrically inconsistent tracks when their predictions are chained across views. We propose MV-RoMa, a multi-view dense matching model that jointly estimates dense correspondences from a source image to multiple co-visible targets. Specifically, we design an efficient model architecture which avoids high computational cost of full cross-attention for multi-view feature interaction: (i) multi-view encoder that leverages pair-wise matching results as a geometric prior, and (ii) multi-view matching refiner that refines correspondences using pixel-wise attention. Additionally, we propose a post-processing strategy that integrates our model's consistent multi-view correspondences as high-quality tracks for SfM. Across diverse and challenging benchmarks, MV-RoMa produces more reliable correspondences and substantially denser, more accurate 3D reconstructions than existing sparse and dense matching methods. Project page: https://icetea-cv.github.io/mv-roma/.

多视角匹配3D重建SfM稠密对应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。