arXiv:2509.12938cs.CV2025-09被引 1

用多视图嵌入袋突破3D高斯点云语义瓶颈

Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings

  • 以物体级嵌入袋替代逐高斯点语义融合
  • 实现开放词汇3D物体检索与2D/3D任务无缝迁移
  • 适合需要精确语义理解的AR/VR与机器人应用

新视角合成近年因3D高斯点云(3DGS)取得显著进展,支持实时逼真渲染。然而,高斯点云固有的模糊性限制了其在AR/VR和机器人中的3D场景理解应用。现有方法通过2D基础模型蒸馏学习语义,但因α混合平均化对象间语义,无法实现3D级理解。本文提出颠覆性方案:完全绕过可微渲染获取语义。核心思路是利用预分解的物体级高斯点,通过多视图CLIP特征聚合构建全面的“嵌入袋”,整体描述物体。该方法实现:(1) 通过文本查询与物体级嵌入对比,实现精准开放词汇物体检索;(2) 无缝任务适配:将物体ID传播至像素实现2D分割,或传播至高斯点实现3D提取。实验表明,本方法有效克服3D开放词汇物体提取挑战,同时在2D开放词汇分割上性能接近最先进水平,代价极小。

原文摘要 · Abstract (English)

Novel view synthesis has seen significant advancements with 3D Gaussian Splatting (3DGS), enabling real-time photorealistic rendering. However, the inherent fuzziness of Gaussian Splatting presents challenges for 3D scene understanding, restricting its broader applications in AR/VR and robotics. While recent works attempt to learn semantics via 2D foundation model distillation, they inherit fundamental limitations: alpha blending averages semantics across objects, making 3D-level understanding impossible. We propose a paradigm-shifting alternative that bypasses differentiable rendering for semantics entirely. Our key insight is to leverage predecomposed object-level Gaussians and represent each object through multiview CLIP feature aggregation, creating comprehensive "bags of embeddings" that holistically describe objects. This allows: (1) accurate open-vocabulary object retrieval by comparing text queries to object-level (not Gaussian-level) embeddings, and (2) seamless task adaptation: propagating object IDs to pixels for 2D segmentation or to Gaussians for 3D extraction. Experiments demonstrate that our method effectively overcomes the challenges of 3D open-vocabulary object extraction while remaining comparable to state-of-the-art performance in 2D open-vocabulary segmentation, ensuring minimal compromise.

3D理解开放词汇高斯点云多视图嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。