通过分层语义图与最优传输规划,提升视觉语言导航的长程决策能力。
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

- 构建动态分层语义场景图,融合多尺度环境表征。
- 基于最优传输理论选择长期目标,保障路径最优性。
- 适合研究视觉语言导航与智能体决策的学者使用。
视觉语言导航在连续环境(VLN-CE)中对自主智能体构成巨大挑战,需无缝融合自然语言指令与视觉观测以导航复杂三维室内空间。现有方法在长程任务中常因场景理解有限、规划效率低及缺乏稳健决策框架而失效。本文提出分层语义增强导航(HSAN)框架,通过三项协同创新重新定义VLN-CE:首先,利用视觉语言模型构建动态分层语义场景图,捕捉从物体到区域再到区域的多层次环境表征,支持细致的空间推理;其次,采用基于最优传输的拓扑规划器,依托坎托罗维奇对偶性,平衡语义相关性与空间可达性,实现理论上最优的目标选择;最后,设计图感知强化学习策略,精确执行子目标并有效避障。通过融合谱图理论、最优传输与先进多模态学习,HSAN克服了传统静态地图与启发式规划的缺陷。在多个具有挑战性的VLN-CE数据集上进行的大量实验表明,该方法实现了领先性能,在导航成功率和未见环境泛化能力上均有显著提升。
原文摘要 · Abstract (English)
Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of natural language instructions and visual observations to navigate complex 3D indoor spaces. Existing approaches often falter in long-horizon tasks due to limited scene understanding, inefficient planning, and lack of robust decision-making frameworks. We introduce the \textbf{Hierarchical Semantic-Augmented Navigation (HSAN)} framework, a groundbreaking approach that redefines VLN-CE through three synergistic innovations. First, HSAN constructs a dynamic hierarchical semantic scene graph, leveraging vision-language models to capture multi-level environmental representations, from objects to regions to zones, enabling nuanced spatial reasoning. Second, it employs an optimal transport-based topological planner, grounded in Kantorovich's duality, to select long-term goals by balancing semantic relevance and spatial accessibility with theoretical guarantees of optimality. Third, a graph-aware reinforcement learning policy ensures precise low-level control, navigating subgoals while robustly avoiding obstacles. By integrating spectral graph theory, optimal transport, and advanced multi-modal learning, HSAN addresses the shortcomings of static maps and heuristic planners prevalent in prior work. Extensive experiments on multiple challenging VLN-CE datasets demonstrate that HSAN achieves state-of-the-art performance, with significant improvements in navigation success and generalization to unseen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。