用隐式方法从无姿态多视角图重建连续3D几何与外观
IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation

- 隐式建模连续几何,通过坐标系内任意点查询生成结果
- 在多个任务上超越现有方法,实现高质量新视角合成与深度估计
- 适合需要高精度3D重建与泛化能力的研究者
从无姿态的多视角图像中重建一致的3D几何与外观是计算机视觉中的基础但具挑战性问题。现有视觉几何基础模型通常通过回归像素对齐的点图来显式预测几何,常存在冗余和几何连续性不足的问题。我们提出IVGT,一种隐式视觉几何变换器,可从无姿态多视角图像中隐式建模连续且一致的几何。该方法在规范坐标系中学习连续的神经场景表示,支持任意3D位置的连续空间查询,利用轻量解码器检索局部特征,预测有符号距离(SDF)值和颜色。它能直接提取连续一致的表面几何,支持从任意视角渲染RGB图像、深度图和法线图。我们通过多数据集联合优化,结合2D监督与3D几何正则化训练IVGT。实验表明,该方法在跨场景泛化方面表现优异,在网格与点云重建、新视角合成、深度与法线估计、相机位姿估计等任务上均取得强性能。
原文摘要 · Abstract (English)
Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundation models predict explicit geometry by regressing pixel-aligned pointmaps, often suffering from redundancy and limited geometric continuity. We propose IVGT, an Implicit Visual Geometry Transformer that implicitly models continuous and coherent geometry from pose-free multi-view images. This formulation learns a continuous neural scene representation in a canonical coordinate system and supports continuous spatial queries at any 3D positions, retrieving local features to predict signed distance (SDF) values and colors using lightweight decoders. It allows direct extraction of continuous and coherent surface geometry, enabling rendering of RGB images, depth maps, and surface normal maps from arbitrary viewpoints. We train IVGT via multi-dataset joint optimization with 2D supervision and 3D geometric regularization. IVGT demonstrates generalization across scenes and achieves strong performance on various tasks, including mesh and point cloud reconstruction, novel view synthesis, depth and surface normal estimation, and camera pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。