用大模型+动态场景图,让机器人更聪明地找东西。
SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation

- 大模型负责全局规划,分解任务为可执行步骤。
- 在i-THOR和RoboTHOR上导航效率提升,零样本泛化能力强。
- 适合做智能机器人导航,尤其复杂环境下的长程任务。
目标导向导航(ObjNav)要求具身智能体仅通过第一人称视觉观测自主定位指定目标。现有单体方法在长程推理上表现不佳,且难以泛化到新环境。为此,我们提出SAGE-Nav,一种融合大语言模型(LLM)与动态场景图的分层框架。关键在于将异步的全局语义规划与高频反应式控制解耦。LLM作为全局规划器,将抽象指令分解为一系列语义锚定的航点;为将计划转化为密集多模态引导,设计了分层场景图编码器(HSGE),利用关系图卷积生成保留语义与空间拓扑结构的嵌入表示;进一步提出目标感知对齐融合网络(GAFN),通过自适应门控机制与显式归纳偏置,动态融合实时感知与结构先验,确保低层策略的鲁棒视觉-拓扑对齐。在i-THOR与RoboTHOR环境中的大量实验表明,SAGE-Nav达到当前最优性能,在导航效率与零样本泛化方面均有显著提升,同时保持物理机器人部署所需的低控制延迟。
原文摘要 · Abstract (English)
Object-Goal Navigation (ObjNav) requires embodied agents to autonomously locate specified targets using only egocentric visual observations. Existing monolithic methods struggle with long-horizon reasoning and generalize poorly to novel environments. To address these limitations, we propose SAGE-Nav, a novel hierarchical framework that integrates the reasoning capabilities of Large Language Models (LLMs) with dynamic scene graphs. Crucially, it decouples asynchronous global semantic planning from the high-frequency reactive control loop. The LLM serves as a global planner, decomposing abstract instructions into a sequence of semantically grounded waypoints. To translate these plans into dense multi-modal guidance, we design a Hierarchical Scene Graph Encoder (HSGE) that leverages relational graph convolutions to produce structure-aware embeddings preserving both semantic and spatial topology. Furthermore, we develop the Goal-aware Alignment-Fusion Network (GAFN) to dynamically fuse real-time perception with these structural priors. Using an adaptive gating mechanism with an explicit inductive bias, GAFN ensures robust visual-topological alignment for the low-level policy. Extensive evaluations in the i-THOR and RoboTHOR environments demonstrate that SAGE-Nav achieves state-of-the-art performance, delivering substantial gains in navigation efficiency and zero-shot generalization while maintaining the low control latency required for physical robotic deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。