arXiv:2601.01872cs.RO2026-01中稿 · ICRA被引 3

用场景图+大模型实现动态户外环境的长期自主导航

CausalNav: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios

  • 构建多层级语义场景图,融合地图与物体信息
  • 支持开放词汇查询,实现在动态环境中的长期稳定导航
  • 适合需要在复杂户外环境长期运行的机器人系统

大规模户外环境中自主语言引导导航仍是移动机器人的重要挑战,主要源于语义推理困难、动态环境干扰及长期稳定性问题。本文提出 CausalNav,首个基于场景图的语义导航框架,专为动态户外环境设计。利用大语言模型构建多层级语义场景图(称作 Embodied Graph),将粗粒度地图数据与细粒度物体实体分层整合。该图作为可检索的知识库,支持检索增强生成(RAG),实现开放词汇下的语义导航与长距离规划。通过融合实时感知与离线地图数据,Embodied Graph 在动态户外环境中支持跨空间粒度的鲁棒导航。动态物体在场景图构建与分层规划模块中均被显式处理。Embodied Graph 在时间窗口内持续更新,反映环境变化,支持实时语义导航。大量仿真与真实世界实验表明,该系统具备优异的鲁棒性与效率。

原文摘要 · Abstract (English)

Autonomous language-guided navigation in large-scale outdoor environments remains a key challenge in mobile robotics, due to difficulties in semantic reasoning, dynamic conditions, and long-term stability. We propose CausalNav, the first scene graph-based semantic navigation framework tailored for dynamic outdoor environments. We construct a multi-level semantic scene graph using LLMs, referred to as the Embodied Graph, that hierarchically integrates coarse-grained map data with fine-grained object entities. The constructed graph serves as a retrievable knowledge base for Retrieval-Augmented Generation (RAG), enabling semantic navigation and long-range planning under open-vocabulary queries. By fusing real-time perception with offline map data, the Embodied Graph supports robust navigation across varying spatial granularities in dynamic outdoor environments. Dynamic objects are explicitly handled in both the scene graph construction and hierarchical planning modules. The Embodied Graph is continuously updated within a temporal window to reflect environmental changes and support real-time semantic navigation. Extensive experiments in both simulation and real-world settings demonstrate superior robustness and efficiency.

自主导航场景图动态环境大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。