arXiv:2508.06990cs.RO2025-08被引 6

用符号化场景图预演未来环境,让智能体更聪明地导航

Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

  • 构建动态分层场景图,结合大模型预测未知区域
  • 在HM3D和HSSD上成功率分别达65.4%和66.8%
  • 适合需要跨房间跨楼层导航的机器人应用

语义导航要求智能体在未见过的环境中定位指定目标。我们提出SGImagineNav,一种创新的想象式导航框架,利用符号化世界建模主动构建全局环境表征。该框架维护一个动态演化的分层场景图,并使用大语言模型预测和探索环境中的未知部分。与仅依赖历史观测的方法不同,这种想象式场景图提供了更丰富的语义上下文,使智能体能够主动估计目标位置。在此基础上,SGImagineNav采用自适应导航策略:当存在语义捷径时加以利用,否则探索未知区域以获取更多信息。该策略持续扩展已知环境并积累有价值的语义上下文,最终引导智能体抵达目标。SGImagineNav在真实场景和仿真基准上进行评估,表现优于以往方法,在HM3D和HSSD上的成功率分别提升至65.4%和66.8%,并在真实环境中实现跨楼层、跨房间导航,验证了其有效性和泛化能力。

原文摘要 · Abstract (English)

Semantic navigation requires an agent to navigate toward a specified target in an unseen environment. Employing an imaginative navigation strategy that predicts future scenes before taking action, can empower the agent to find target faster. Inspired by this idea, we propose SGImagineNav, a novel imaginative navigation framework that leverages symbolic world modeling to proactively build a global environmental representation. SGImagineNav maintains an evolving hierarchical scene graphs and uses large language models to predict and explore unseen parts of the environment. While existing methods solely relying on past observations, this imaginative scene graph provides richer semantic context, enabling the agent to proactively estimate target locations. Building upon this, SGImagineNav adopts an adaptive navigation strategy that exploits semantic shortcuts when promising and explores unknown areas otherwise to gather additional context. This strategy continuously expands the known environment and accumulates valuable semantic contexts, ultimately guiding the agent toward the target. SGImagineNav is evaluated in both real-world scenarios and simulation benchmarks. SGImagineNav consistently outperforms previous methods, improving success rate to 65.4 and 66.8 on HM3D and HSSD, and demonstrating cross-floor and cross-room navigation in real-world environments, underscoring its effectiveness and generalizability.

智能体导航场景图大模型语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。