arXiv:2605.26500cs.CV2026-05ICCV被引 12

用3D高斯表示场景,结合语义分组提升视觉语言导航能力

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

论文配图:3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
图 1 · 摘自论文原文
  • 基于伪激光点云初始化可微3D高斯,构建实时场景地图
  • 通过开放世界语义分组将高斯聚类为物体或场景类别
  • 多粒度空间语义融合策略,提升复杂环境导航表现

视觉语言导航(VLN)要求智能体根据自然语言指令在复杂3D环境中移动,需具备全面的场景理解能力。现有方法虽引入多种场景表征以增强空间感知,但常忽略复杂3D几何与丰富语义,限制了在多样及未见环境中的泛化能力。本文提出一种3D高斯地图,将环境表示为可微3D高斯集合,并设计相应的VLN导航策略。具体地,通过稀疏伪激光点云在线构建本体场景地图,为场景理解提供几何先验;每个高斯原语进一步通过开放世界语义分组操作进行丰富,依据其在开放世界中属于物体实例或背景类别进行聚类,形成统一的3D高斯地图。在此基础上,设计多层级动作预测策略,融合多粒度空间-语义线索以辅助决策。在R2R、R4R和REVERIE三个公开基准上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene representations to enhance spatial awareness, they often neglect the complex 3D geometry and rich semantics in VLN scenarios, limiting the ability to generalize across diverse and unseen environments. To address these challenges, this work proposes a 3D Gaussian Map that represents the environment as a set of differentiable 3D Gaussians and accordingly develops a navigation strategy for VLN. Specifically, Egocentric Scene Map is constructed online by initializing 3D Gaussians from sparse pseudo-lidar point clouds, providing informative geometric priors for scene understanding. Each Gaussian primitive is further enriched through Open-Set Semantic Grouping operation, which groups 3D Gaussians based on their membership in object instances or stuff categories within the open world, resulting in a unified 3D Gaussian Map. Building on this map, Multi-Level Action Prediction strategy, which combines spatial-semantic cues at multiple granularities, is designed to assist agents in decision-making. Extensive experiments conducted on three public benchmarks (i.e., R2R, R4R, and REVERIE) validate the effectiveness of our method.

视觉导航3D高斯语义分组多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。