arXiv:2412.01807cs.CV2024-12被引 18

用简单方法提升3D场景语言理解效率,速度提升100倍

Occam's LGS: An Efficient Approach for Language Gaussian Splatting

  • 基于概率建模简化多视角语言特征融合流程
  • 在不压缩的前提下实现比之前快100倍的训练速度
  • 适合需要快速交互式3D场景编辑的研究者

高斯点阵(Gaussian Splatting)是广泛采用的3D场景表示方法,具备高效且高质量的重建与渲染能力。其成功关键在于使用高斯分布集合表示场景,具有可解释性和可适应性。为超越视觉表示,近期研究将语义视觉-语言特征引入高斯点阵,支持开放集任务。然而,现有方法依赖多视图语言特征聚合,流程复杂,计算开销大、训练时间长。本文表明,复杂的语言3D高斯点阵管道实属冗余。我们基于语言高斯点阵的概率形式,应用奥卡姆剃刀原则,提出一种高效的加权多视图特征聚合方法。该方法在不进行任何压缩的情况下,实现两数量级的速度提升,并达到当前最优性能,支持便捷的场景操作。项目页面:https://insait-institute.github.io/OccamLGS/

原文摘要 · Abstract (English)

TL;DR: Gaussian Splatting is a widely adopted approach for 3D scene representation, offering efficient, high-quality reconstruction and rendering. A key reason for its success is the simplicity of representing scenes with sets of Gaussians, making it interpretable and adaptable. To enhance understanding beyond visual representation, recent approaches extend Gaussian Splatting with semantic vision-language features, enabling open-set tasks. Typically, these language features are aggregated from multiple 2D views, however, existing methods rely on cumbersome techniques, resulting in high computational costs and longer training times. In this work, we show that the complicated pipelines for language 3D Gaussian Splatting are simply unnecessary. Instead, we follow a probabilistic formulation of Language Gaussian Splatting and apply Occam's razor to the task at hand, leading to a highly efficient weighted multi-view feature aggregation technique. Doing so offers us state-of-the-art results with a speed-up of two orders of magnitude without any compression, allowing for easy scene manipulation. Project Page: https://insait-institute.github.io/OccamLGS/

3D生成高斯点阵语言模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。