用语言和几何联合引导的稀疏体素,统一建模3D场景的外观、语义与结构。
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
- 以稀疏体素为单元,构建外观、密度、特征和置信度四场统一表示
- 通过特征调制与深度相关正则化,实现语言与几何知识的有效迁移
- 适合需要高精度三维理解与重建的视觉任务,如机器人导航
现有3D开放词汇场景理解方法多侧重将2D基础模型的语言特征迁移到3D特征场,却忽视了场景外观、语义与几何之间的协同关系。这导致理解结果偏离真实几何结构,且与重建过程脱节。本文提出一种新方法:利用语言与几何共同引导的稀疏体素表示,在统一框架中全面建模3D场景的外观、语义与几何。具体地,采用3D稀疏体素作为基本单元,引入外观场、密度场、特征场与置信度场进行综合表征。为增强各场间的协同性,设计特征调制模块,并从2D基础模型中蒸馏语言特征至3D模型。同时,通过深度相关正则化与模式一致性正则化,将几何基础模型中的几何知识迁移到特征场中。这些组件协同作用,显著提升对场景整体结构的理解与重建能力。大量实验表明,该方法在整体场景理解与重建任务上优于当前最优方法。
原文摘要 · Abstract (English)
Existing 3D open-vocabulary scene understanding methods mostly emphasize distilling language features from 2D foundation models into 3D feature fields, but largely overlook the synergy among scene appearance, semantics, and geometry. As a result, scene understanding often deviates from the underlying geometric structure of scenes and becomes decoupled from the reconstruction process. In this work, we propose a novel approach that leverages language and geometry grounded sparse voxel representations to comprehensively model appearance, semantics, and geometry within a unified framework. Specifically, we use 3D sparse voxels as primitives and employ an appearance field, a density field, a feature field, and a confidence field to holistically represent a 3D scene. To promote synergy among the appearance, density, and feature fields, we construct a feature modulation module and distill language features from a 2D foundation model into our 3D scene model. In addition, we integrate geometric distillation into feature field distillation to transfer geometric knowledge from a geometry foundation model to our 3D scene representations via depth correlation regularization and pattern consistency regularization. These components work together to synergistically model the appearance, semantics, and geometry of the 3D scene within a unified framework. Extensive experiments demonstrate that our approach achieves superior overall performance compared with state-of-the-art methods in holistic scene understanding and reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。