不依赖像素对齐,用全局视角重建完整3D场景。
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- 用场景令牌聚合多视角信息,摆脱像素对齐限制。
- 可恢复可见与不可见点,重叠区域结构更合理。
- 适合需要完整几何结构的3D重建任务。
我们提出NOVA3R,一种从无姿态图像集合中进行前馈式非像素对齐3D重建的有效方法。不同于将几何与每条射线预测绑定的像素对齐方法,我们的模型学习全局、视角无关的场景表示,使重建脱离像素对齐约束。该方法解决了像素对齐3D重建的两大局限:(1) 能够通过完整的场景表示恢复可见与不可见点;(2) 在重叠区域生成更物理合理的几何结构,减少重复结构。为此,我们引入场景令牌机制以跨无姿态图像聚合信息,并设计基于扩散的3D解码器,重建完整的非像素对齐点云。在场景级与物体级数据集上的大量实验表明,NOVA3R在重建精度和完整性上均优于现有最先进方法。
原文摘要 · Abstract (English)
We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulation learns a global, view-agnostic scene representation that decouples reconstruction from pixel alignment. This addresses two key limitations in pixel-aligned 3D: (1) it recovers both visible and invisible points with a complete scene representation, and (2) it produces physically plausible geometry with fewer duplicated structures in overlapping regions. To achieve this, we introduce a scene-token mechanism that aggregates information across unposed images and a diffusion-based 3D decoder that reconstructs complete, non-pixel-aligned point clouds. Extensive experiments on both scene-level and object-level datasets demonstrate that NOVA3R outperforms state-of-the-art methods in terms of reconstruction accuracy and completeness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。