arXiv:2512.04012cs.CV2025-12被引 17

无需额外训练,模型自动生成去噪能力,自动剔除干扰图像。

Emergent Outlier View Rejection in Visual Geometry Grounded Transformers

  • 发现模型特定层能自发抑制干扰图像,无需显式设计
  • 在含20%-80%合成干扰图的测试中保持稳定性能
  • 适用于真实场景下无监督的3D重建,无需额外标注

从真实世界图像集合中可靠进行3D重建常受'噪声'图像干扰——即与其它图像视图重叠极少的无关输入。传统SfM流水线通过几何验证和异常值剔除处理此类问题,但前馈式3D重建模型缺乏此类显式机制,导致在真实场景下性能下降。本文发现,现有前馈重建模型(如VGGT)虽未设计异常值剔除机制或噪声感知训练,仍能内在区分干扰图像。通过在不同比例合成干扰图像下的深入分析,我们识别出一个具有自然异常值抑制行为的特定层。进一步探查表明,该层编码了可区分的内部表征,具备有效的噪声过滤能力。我们仅利用此特性,在不增加微调或监督的前提下,实现前馈3D重建中的异常视图剔除。在受控及真实数据集上的大量实验表明,该隐式过滤机制具有一致性,并在多种场景中表现出良好泛化能力。

原文摘要 · Abstract (English)

Reliable 3D reconstruction from in-the-wild image collections is often hindered by "noisy" images-irrelevant inputs with little or no view overlap with others. While traditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed-forward 3D reconstruction models lack these explicit mechanisms, leading to degraded performance under in-the-wild conditions. In this paper, we discover that the existing feed-forward reconstruction model, e.g., VGGT, despite lacking explicit outlier-rejection mechanisms or noise-aware training, can inherently distinguish distractor images. Through an in-depth analysis under varying proportions of synthetic distractors, we identify a specific layer that naturally exhibits outlier-suppressing behavior. Further probing reveals that this layer encodes discriminative internal representations that enable an effective noise-filtering capability, which we simply leverage to perform outlier-view rejection in feed-forward 3D reconstruction without any additional fine-tuning or supervision. Extensive experiments on both controlled and in-the-wild datasets demonstrate that this implicit filtering mechanism is consistent and generalizes well across diverse scenarios.

3D重建去噪自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。