arXiv:2511.22686cs.CV2025-11被引 2

3D基础模型能自发理解极端视角几何,无需专门训练。

Emergent Extreme-View Geometry in 3D Foundation Models

  • 通过微调少量偏置项,提升模型在极端视角下的三维推理能力。
  • 在未见过的场景上,相对位姿估计精度显著提升,且不影响单图深度和点云质量。
  • 新构建了MegaUnScene基准,适合评估模型在极端视角下的泛化能力。

3D基础模型(3DFMs)近期推动了3D视觉的发展,可直接从图像中联合预测深度、位姿和点云。然而,它们在极端非重叠视角下的推理能力仍待探索。本文研究其内部表征,发现3DFMs虽未针对此类条件训练,却表现出对极端视角几何的自发理解。为此,我们提出一种轻量级对齐方案,仅微调骨干网络中少量偏置项,冻结所有解码头,显著提升极端视角下的相对位姿估计性能,同时不降低单图深度与点云质量。此外,我们构建了MegaUnScene,一个现有3DFMs未见过的互联网场景新基准,包含专用于相对位姿估计与密集3D重建的测试集。全部代码与数据将公开。

原文摘要 · Abstract (English)

3D foundation models (3DFMs) have recently transformed 3D vision, enabling joint prediction of depths, poses, and point maps directly from images. Yet their ability to reason under extreme, non-overlapping views remains largely unexplored. In this work, we study their internal representations and find that 3DFMs exhibit an emergent understanding of extreme-view geometry, despite never being trained for such conditions. To further enhance these capabilities, we introduce a lightweight alignment scheme that refines their internal 3D representation by tuning only a small subset of backbone bias terms, leaving all decoder heads frozen. This targeted adaptation substantially improves relative pose estimation under extreme viewpoints without degrading per-image depth or point quality. Additionally, we contribute MegaUnScene, a new benchmark of Internet scenes unseen by existing 3DFMs, with dedicated test splits for both relative pose estimation and dense 3D reconstruction. All code and data will be released.

3D生成视觉推理基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。