arXiv:2604.01388cs.CV2026-04被引 1

用稀疏体素融合语言特征,提升3D场景理解的精度与细节。

LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding

论文配图:LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
图 1 · 摘自论文原文
  • 采用稀疏体素栅格化构建结构化几何基础,避免高斯点云重叠问题。
  • 在点云理解任务上达顶尖性能,2D/3D物体检索也表现优异。
  • 适合关注开放词汇3D理解与视觉语言对齐的研究者。

当前开放词汇3D场景理解依赖3D高斯点云(3DGS)将视觉-语言特征映射至3D空间。然而我们发现两类关键局限:无序重叠的高斯点导致空间模糊性,需依赖概率注册;基于物体级掩码池化引发多层级语义模糊,削弱细粒度信息。为此,本文提出新框架,采用稀疏体素栅格化(SVRaster)作为结构化、互斥的几何表示。通过引入单目深度与法向先验对SVRaster进行正则化,建立稳定几何基底,实现确定性、置信度感知的特征注册,抑制3DGS中常见的语义泄漏现象。此外,借助AM-RADIO基础模型的密集对齐特性,有效缓解多层级模糊,无需分层训练的计算开销。本方法在开放词汇点云理解任务上达到当前最优性能,在3D与2D物体检索基准上亦表现强劲。

原文摘要 · Abstract (English)

Recent advancements in open-vocabulary 3D scene understanding heavily rely on 3D Gaussian Splatting (3DGS) to register vision-language features into 3D space. However, we identify two critical limitations in these approaches: the spatial ambiguity arising from unstructured, overlapping Gaussians which necessitates probabilistic feature registration, and the multi-level semantic ambiguity caused by pooling features over object-level masks, which dilutes fine-grained details. To address these challenges, we present a novel framework that leverages Sparse Voxel Rasterization (SVRaster) as a structured, disjoint geometry representation. By regularizing SVRaster with monocular depth and normal priors, we establish a stable geometric foundation. This enables a deterministic, confidence-aware feature registration process and suppresses the semantic bleeding artifact common in 3DGS. Furthermore, we resolve multi-level ambiguity by exploiting the emerging dense alignment properties of the AM-RADIO foundation model, avoiding the computational overhead of hierarchical training methods. Our approach achieves state-of-the-art performance on Open Vocabulary Point Cloud Understanding, and highly competitive results on 3D and 2D Object Retrieval benchmarks.

3D理解视觉语言稀疏体素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。