arXiv:2412.05826cs.CV2024-12CVPR被引 19

用3D几何特征提升视觉歧义识别,让3D重建更准

Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features

  • 基于MASt3R的3D感知特征,用Transformer增强歧义检测
  • 在跨场景测试中召回率和准确率均显著提升,优于旧方法
  • 可无缝接入主流3D重建流程,适合复杂真实场景

精确的3D重建常受视觉相似但结构不同的表面(即伪孪生体)干扰,导致误匹配,扭曲结构光恢复(SfM)过程并降低精度。以往方法依赖在精选数据集上训练的CNN分类器,但在多样化真实场景中泛化能力差且需大量调参。本文提出Doppelgangers++,通过引入包含日常场景地理标签图像的多样化训练数据集,提升模型鲁棒性;进一步设计基于Transformer的分类器,利用MASt3R模型提取的3D感知特征,在域内与域外测试中均实现更高精度与召回率。该方法可无缝集成至标准SfM及MASt3R-SfM流程,兼具高效与适应性。为评估3D重建质量,我们提出一种基于地理标签的自动化验证方法,无需人工检查。大量实验证明,Doppelgangers++显著改善了配对视觉歧义消除效果,并提升了复杂多变场景下的3D重建质量。

原文摘要 · Abstract (English)

Accurate 3D reconstruction is frequently hindered by visual aliasing, where visually similar but distinct surfaces (aka, doppelgangers), are incorrectly matched. These spurious matches distort the structure-from-motion (SfM) process, leading to misplaced model elements and reduced accuracy. Prior efforts addressed this with CNN classifiers trained on curated datasets, but these approaches struggle to generalize across diverse real-world scenes and can require extensive parameter tuning. In this work, we present Doppelgangers++, a method to enhance doppelganger detection and improve 3D reconstruction accuracy. Our contributions include a diversified training dataset that incorporates geo-tagged images from everyday scenes to expand robustness beyond landmark-based datasets. We further propose a Transformer-based classifier that leverages 3D-aware features from the MASt3R model, achieving superior precision and recall across both in-domain and out-of-domain tests. Doppelgangers++ integrates seamlessly into standard SfM and MASt3R-SfM pipelines, offering efficiency and adaptability across varied scenes. To evaluate SfM accuracy, we introduce an automated, geotag-based method for validating reconstructed models, eliminating the need for manual inspection. Through extensive experiments, we demonstrate that Doppelgangers++ significantly enhances pairwise visual disambiguation and improves 3D reconstruction quality in complex and diverse scenarios.

3D重建视觉歧义TransformerMASt3R

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。