arXiv:2412.08331cs.CV2024-12中稿 · ACM MM 2025被引 17

用少量视角快速构建带语义的3D场景,适合实时应用。

SLGaussian: Fast Language Gaussian Splatting in Sparse Views

  • 通过视频追踪保证分割一致,用低维索引嵌入语言特征。
  • 两视图下在LERF和3D-OVS数据集上性能领先,mIoU更高。
  • 单场景推理<30秒,开放词汇查询仅需0.011秒/次。

3D语义场学习对自动驾驶、AR/VR和机器人等应用至关重要,需从有限视角准确理解3D场景。现有方法在稀疏视图下表现不佳,依赖低效的逐场景多视图优化,难以应用于实际任务。为此,我们提出SLGaussian,一种从前向视角构建3D语义场的方法,可直接生成基于3DGS的场景。通过视频追踪确保SAM分割一致性,并使用低维索引编码高维CLIP特征,高效将语言信息嵌入3D空间。在两视图稀疏3D物体查询与分割任务中,于LERF和3D-OVS数据集上,SLGaussian在选定的IoU、定位精度和mIoU指标上优于现有方法。此外,模型实现场景推理时间低于30秒,开放词汇查询速度达每查询0.011秒。

原文摘要 · Abstract (English)

3D semantic field learning is crucial for applications like autonomous navigation, AR/VR, and robotics, where accurate comprehension of 3D scenes from limited viewpoints is essential. Existing methods struggle under sparse view conditions, relying on inefficient per-scene multi-view optimizations, which are impractical for many real-world tasks. To address this, we propose SLGaussian, a feed-forward method for constructing 3D semantic fields from sparse viewpoints, allowing direct inference of 3DGS-based scenes. By ensuring consistent SAM segmentations through video tracking and using low-dimensional indexing for high-dimensional CLIP features, SLGaussian efficiently embeds language information in 3D space, offering a robust solution for accurate 3D scene understanding under sparse view conditions. In experiments on two-view sparse 3D object querying and segmentation in the LERF and 3D-OVS datasets, SLGaussian outperforms existing methods in chosen IoU, Localization Accuracy, and mIoU. Moreover, our model achieves scene inference in under 30 seconds and open-vocabulary querying in just 0.011 seconds per query.

3D语义场稀疏视图快速推理开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。