用3D高斯图增强场景理解,提升物体定位与空间关系推理能力
GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
- 通过动态聚类策略避免特征压缩,提高物体分割精度
- 融合2D模型提取的属性与空间关系,构建更完整的场景图
- 引入3D一致性验证模块,过滤不合理空间关系,适合复杂场景交互
3D高斯点阵(3DGS)的进步显著提升了语义场景理解能力,支持自然语言查询定位场景中的物体。然而,现有方法主要依赖将压缩后的CLIP特征嵌入3D高斯,导致物体分割精度低且缺乏空间推理能力。为此,我们提出GaussianGraph,一种通过自适应语义聚类和场景图生成增强3DGS场景理解的新框架。我们设计了'Control-Follow'聚类策略,动态适应场景尺度与特征分布,避免特征压缩,显著提升分割准确率。同时,通过2D基础模型提取对象属性与空间关系,丰富场景表征。针对空间关系不准确问题,提出3D校正模块,基于空间一致性验证过滤不合理关系,确保可靠场景图构建。在三个数据集上的大量实验表明,GaussianGraph在语义分割和物体定位任务上均优于当前最优方法,为复杂场景理解与交互提供了鲁棒解决方案。
原文摘要 · Abstract (English)
Recent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low object segmentation accuracy and lack spatial reasoning capabilities. To address these limitations, we propose GaussianGraph, a novel framework that enhances 3DGS-based scene understanding by integrating adaptive semantic clustering and scene graph generation. We introduce a "Control-Follow" clustering strategy, which dynamically adapts to scene scale and feature distribution, avoiding feature compression and significantly improving segmentation accuracy. Additionally, we enrich scene representation by integrating object attributes and spatial relations extracted from 2D foundation models. To address inaccuracies in spatial relationships, we propose 3D correction modules that filter implausible relations through spatial consistency verification, ensuring reliable scene graph construction. Extensive experiments on three datasets demonstrate that GaussianGraph outperforms state-of-the-art methods in both semantic segmentation and object grounding tasks, providing a robust solution for complex scene understanding and interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。