arXiv:2605.29879cs.CVcs.RO2026-05

构建动态3D高斯场景图,实现长期场景理解与目标定位

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

论文配图:DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
图 1 · 摘自论文原文
  • 融合概率体素与3D高斯,实现跨模态实例关联与增量语义映射
  • 在自重建地图上零样本3D视觉接地任务中性能领先,达85.3%准确率
  • 适用于机器人长期导航与动态环境更新,支持多模态推理

将开放词汇语义信息融入动态3D场景表示对于长期具身场景理解至关重要。现有方法常因视图间线索不完整导致实例关联脆弱,且难以处理物体层面的拓扑变化,限制了长期机器人任务执行。当前方法或依赖简单特征匹配而缺乏显式空间推理,或假设存在离线真值3D几何。为此,我们提出DGSG-Mind,一种结合具身推理代理的混合实例感知3D高斯动态场景图系统。该系统通过概率体素网格与显式3D高斯耦合,实现鲁棒的跨模态实例融合与增量语义映射;通过基于高斯的视觉重定位与由几何-语义一致性引导的局部掩码精炼,应对动态变化。在实例高斯地图基础上,DGSG-Mind进一步构建分层场景图,并开发3D高斯心智(3D Gaussian Mind),整合结构关系、时空语义信息及视觉标注的感兴趣区域高斯渲染,支持多模态推理。大量实验表明,DGSG-Mind在自重建地图上零样本3D视觉接地任务中表现最佳,准确率达85.3%;同时在3D开放词汇语义分割与场景重建任务中也表现出色。我们进一步在真实机器人上部署,验证其目标导向推理与动态更新能力。项目页面见:https://icr-lab.github.io/DGSG-Mind

原文摘要 · Abstract (English)

Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However, existing methods often suffer from fragile instance association due to incomplete cross-view cues, while their limited ability to handle object-level topological changes restricts long-term robotic task execution. Moreover, current 3D scene understanding methods either rely on simple feature matching without explicit spatial reasoning or assume offline ground-truth 3D geometry. To address these challenges, we present DGSG-Mind, a hybrid instance-aware 3D Gaussian dynamic scene graph system with an embodied reasoning agent. Our system couples a probabilistic voxel grid with explicit 3D Gaussians to enable robust cross-modal instance fusion and incremental semantic mapping. It handles dynamic changes through Gaussian-based visual relocalization and localized masked refinement guided by geometric-semantic consistency. Built on the instance Gaussian map, DGSG-Mind further constructs a hierarchical scene graph and develops the 3D Gaussian Mind, which integrates structural relations, spatial-semantic information, and visually annotated RoI Gaussian renderings for multimodal reasoning. Extensive experiments show that DGSG-Mind achieves the best zero-shot 3DVG performance among methods operating on self-reconstructed maps, while also delivering strong performance in 3D open-vocabulary semantic segmentation and scene reconstruction. We further deploy DGSG-Mind on real-world robots to demonstrate its target-oriented reasoning and dynamic update capabilities. The project page of DGSG-Mind is available at https://icr-lab.github.io/DGSG-Mind

3D高斯场景图机器人动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。