用自生成真值优化3D特征对齐,提升图像重建与相机定位精度
Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- 利用模型自身输出构建伪真值,通过投影一致性损失训练轻量适配器
- 在NVS和相机位姿估计上达到当前最优性能,显著提升多视角几何一致性
- 适合需要高保真3D重建与鲁棒姿态估计的视觉任务研究者
传统新视角合成依赖带有显式3D先验的模型及已知相机参数。近期视觉基础模型如VGGT采用不同路径——通过训练数据和损失目标隐式学习3D知识,直接从前置未标定图像中预测相机参数与3D表示。虽具灵活性,但其特征缺乏显式多视图几何一致性。本文提出Selfi,一种通过特征对齐实现自我优化的3D重建流水线,将VGGT主干转化为高保真3D重建引擎。具体而言,利用模型自身输出作为伪真值,训练轻量级特征适配器,基于重投影一致性损失,将特征映射至新的几何对齐空间,捕捉三维空间邻近性。该方法在新视角合成与相机位姿估计任务上均取得当前最优表现,证明特征对齐对下游3D推理具有显著提升作用。
原文摘要 · Abstract (English)
Novel View Synthesis (NVS) has traditionally relied on models with explicit 3D inductive biases combined with known camera parameters from Structure-from-Motion (SfM) beforehand. Recent vision foundation models like VGGT take an orthogonal approach -- 3D knowledge is gained implicitly through training data and loss objectives, enabling feed-forward prediction of both camera parameters and 3D representations directly from a set of uncalibrated images. While flexible, VGGT features lack explicit multi-view geometric consistency, and we find that improving such 3D feature consistency benefits both NVS and pose estimation tasks. We introduce Selfi, a self-improving 3D reconstruction pipeline via feature alignment, transforming a VGGT backbone into a high-fidelity 3D reconstruction engine by leveraging its own outputs as pseudo-ground-truth. Specifically, we train a lightweight feature adapter using a reprojection-based consistency loss, which distills VGGT outputs into a new geometrically-aligned feature space that captures spatial proximity in 3D. This enables state-of-the-art performance in both NVS and camera pose estimation, demonstrating that feature alignment is a highly beneficial step for downstream 3D reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。