arXiv:2502.16652cs.CV2025-02CVPR被引 59

用语言嵌入直接关联3D高斯点,实现开放词汇场景理解

Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration

  • 直接将CLIP语言嵌入注册到3D高斯点上,跳过渲染过程
  • 在3D语义分割、目标定位等任务上显著超越现有方法
  • 结合产品量化压缩嵌入,无需每场景优化,适合快速部署

我们提出Dr. Splat,一种基于3D高斯溅射的开放词汇3D场景理解新方法。与依赖渲染过程的现有语言嵌入3DGS方法不同,本方法直接将语言对齐的CLIP嵌入分配给每个像素-射线相交的主导高斯点,实现全局3D场景理解。核心是语言特征注册技术,通过该技术将CLIP嵌入映射至对应的3D高斯点。此外,我们采用在通用大规模图像数据上训练的产品量化(PQ)方法,以紧凑方式表示嵌入,无需针对每个场景进行优化。实验表明,该方法在3D感知基准测试中表现优异,包括开放词汇3D语义分割、3D目标定位和3D目标选择任务。视频演示请访问:https://drsplat.github.io/

原文摘要 · Abstract (English)

We introduce Dr. Splat, a novel approach for open-vocabulary 3D scene understanding leveraging 3D Gaussian Splatting. Unlike existing language-embedded 3DGS methods, which rely on a rendering process, our method directly associates language-aligned CLIP embeddings with 3D Gaussians for holistic 3D scene understanding. The key of our method is a language feature registration technique where CLIP embeddings are assigned to the dominant Gaussians intersected by each pixel-ray. Moreover, we integrate Product Quantization (PQ) trained on general large-scale image data to compactly represent embeddings without per-scene optimization. Experiments demonstrate that our approach significantly outperforms existing approaches in 3D perception benchmarks, such as open-vocabulary 3D semantic segmentation, 3D object localization, and 3D object selection tasks. For video results, please visit : https://drsplat.github.io/

3D高斯开放词汇语言嵌入场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。