无需标注的3D重建框架,统一几何与语义理解。
FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views
- 纯前馈架构,仅靠渲染监督实现多视角3D重建。
- 在ScanNet和DL3DV-10K上实现更优的新视角合成与语义分割。
- 适合需要空间与语义联合理解的智能体系统开发。
视觉基础模型的进展已彻底改变几何重建与语义理解。然而,现有方法大多将二者孤立处理,导致冗余流程与误差累积。本文提出FF3R,一种完全无标注的前馈框架,可从任意视角图像序列中统一进行几何与语义推理。不同于以往方法,FF3R无需相机位姿、深度图或语义标签,仅依赖RGB与特征图的渲染监督,建立了一种可扩展的统一3D推理范式。针对前馈特征重建中的两大挑战——全局语义不一致与局部结构不一致,我们提出两项创新:(i) Token-wise Fusion Module通过交叉注意力将语义上下文融入几何标记;(ii) Semantic-Geometry Mutual Boosting机制结合几何引导特征对齐实现全局一致性,以及语义感知体素化确保局部连贯性。在ScanNet与DL3DV-10K上的大量实验表明,FF3R在新视角合成、开放词汇语义分割与深度估计任务中表现卓越,且在真实场景下具有强泛化能力,为需要空间与语义双重理解的具身智能系统铺平道路。
原文摘要 · Abstract (English)
Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches treat these capabilities in isolation, leading to redundant pipelines and compounded errors. This paper introduces FF3R, a fully annotation-free feed-forward framework that unifies geometric and semantic reasoning from unconstrained multi-view image sequences. Unlike previous methods, FF3R does not require camera poses, depth maps, or semantic labels, relying solely on rendering supervision for RGB and feature maps, establishing a scalable paradigm for unified 3D reasoning. In addition, we address two critical challenges in feedforward feature reconstruction pipelines, namely global semantic inconsistency and local structural inconsistency, through two key innovations: (i) a Token-wise Fusion Module that enriches geometry tokens with semantic context via cross-attention, and (ii) a Semantic-Geometry Mutual Boosting mechanism combining geometry-guided feature warping for global consistency with semantic-aware voxelization for local coherence. Extensive experiments on ScanNet and DL3DV-10K demonstrate FF3R's superior performance in novel-view synthesis, open-vocabulary semantic segmentation, and depth estimation, with strong generalization to in-the-wild scenarios, paving the way for embodied intelligence systems that demand both spatial and semantic understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。