发现前馈3D重建模型隐含几何原理,非仅依赖数据统计规律。
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
- 通过注意力机制分析,发现中间层存在本质的对极几何结构。
- 在真实与合成数据上验证,几何理解与对应匹配模式强相关。
- 适合研究视觉几何基础、模型可解释性及多视角重建的学者。
基于Transformer的前馈3D重建模型(如DUSt3R、VGGT和Depth Anything 3)能在单次前向传播中推断相机姿态与密集场景结构。这些模型在大规模监督下训练,引发核心问题:它们是否继承了传统多视图方法的几何原理,还是主要依赖大规模训练带来的学习先验?我们发现,所有三个模型的中间层均涌现出对极几何,并且其与注意力头中的对应匹配模式存在因果关联。为此,我们在三个真实世界数据集和一个受控合成数据集上系统分析其内部表征。通过探测中间特征、分析注意力模式以识别对应匹配,以及在注意力层面进行定向干预,量化几何理解程度。此外,通过施加遮挡、场景模糊及不同相机配置等挑战性输入扰动,评估学习先验的作用,并与经典多阶段重建流水线对比。
原文摘要 · Abstract (English)
Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised fashion, they raise a central question: do these models build upon geometric principles akin to traditional multi-view pipelines, or do they primarily rely on learned priors arising from the large-scale training setup? We find that epipolar geometry emerges within the intermediate layers of all three models and is causally linked to correspondence patterns in attention heads. To study this, we perform a systematic analysis of their internal representations across three real-world datasets and a controlled synthetic dataset. We quantify geometric understanding by probing intermediate features, analyzing attention patterns to identify correspondence matching patterns, and performing targeted interventions at the attention level. Further, we assess the role of learned priors by applying challenging input-level perturbations, such as occlusions, scene ambiguities, and varying camera configurations, and compare them against classical multi-stage reconstruction pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。