让3D场景理解语言,实现自然语言操控大场景。
Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes
- 在3D高斯中嵌入极低维语义瓶颈特征,提升效率。
- 新模块与正则化解决2D语义特征错位问题。
- 适用于需要语言交互的大规模3D场景应用。
在3D表示中嵌入语言场可将几何结构与描述性语义关联,增强空间环境的语义理解,支持通过自然语言查询或编辑场景,有望提升场景检索、导航和多模态推理等任务。然而,面对大规模互联网数据时,现有特征蒸馏方法因语义特征错位及内存与运行效率不足而难以有效学习。为此,我们提出新方法:首先,在底层3D高斯表示中引入极低维语义瓶颈特征,并通过多分辨率、基于特征的哈希编码器处理,显著降低运行时与显存占用;其次,设计衰减下采样模块并提出多种正则化策略,缓解真实2D特征的语义错位问题。我们在野外场景数据集HolyScenes上评估,结果表明该方法在性能与效率上均优于现有方法。
原文摘要 · Abstract (English)
Embedding a language field in a 3D representation enables richer semantic understanding of spatial environments by linking geometry with descriptive meaning. This allows for a more intuitive human-computer interaction, enabling querying or editing scenes using natural language, and could potentially improve tasks like scene retrieval, navigation, and multimodal reasoning. While such capabilities could be transformative, in particular for large-scale scenes, we find that recent feature distillation approaches cannot effectively learn over massive Internet data due to challenges in semantic feature misalignment and inefficiency in memory and runtime. To this end, we propose a novel approach to address these challenges. First, we introduce extremely low-dimensional semantic bottleneck features as part of the underlying 3D Gaussian representation. These are processed by rendering and passing them through a multi-resolution, feature-based, hash encoder. This significantly improves efficiency both in runtime and GPU memory. Second, we introduce an Attenuated Downsampler module and propose several regularizations addressing the semantic misalignment of ground truth 2D features. We evaluate our method on the in-the-wild HolyScenes dataset and demonstrate that it surpasses existing approaches in both performance and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。