通过时空迭代增强,提升视觉语言导航中复杂环境下的感知与决策能力。
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
- 构建分层时空编码与多粒度对齐融合机制,统一全局拓扑与局部网格信息。
- 在多个连续导航数据集上超越现有方法,尤其在干扰环境下表现显著提升。
- 适合研究视觉语言导航、强化学习与多模态感知的学者参考。
在连续环境中的视觉-语言导航(VLN-CE)要求智能体基于自然语言指令在未知连续空间中导航。相比离散场景,其面临两大感知挑战:一是缺乏预定义观测点,导致视觉记忆异质且全局空间相关性弱;二是三维场景中累积的重建误差引入结构噪声,影响局部特征感知。为此,本文提出ST-Booster,一种迭代式时空感知增强框架,通过多粒度感知与指令感知推理提升导航性能。该框架包含三个核心模块:分层时空编码(HSTE)利用拓扑图建模长期全局记忆,以网格地图捕捉短期局部细节;多粒度对齐融合(MGAF)通过几何感知知识融合对齐双地图表示,并经预训练任务迭代优化;价值引导路径点生成(VGWG)生成指导注意力热图(GAH),显式建模环境与指令的相关性并优化路径点选择。大量对比实验与分析表明,ST-Booster在复杂、易受干扰环境中显著优于现有最先进方法。
原文摘要 · Abstract (English)
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to navigate unknown, continuous spaces based on natural language instructions. Compared to discrete settings, VLN-CE poses two core perception challenges. First, the absence of predefined observation points leads to heterogeneous visual memories and weakened global spatial correlations. Second, cumulative reconstruction errors in three-dimensional scenes introduce structural noise, impairing local feature perception. To address these challenges, this paper proposes ST-Booster, an iterative spatiotemporal booster that enhances navigation performance through multi-granularity perception and instruction-aware reasoning. ST-Booster consists of three key modules -- Hierarchical SpatioTemporal Encoding (HSTE), Multi-Granularity Aligned Fusion (MGAF), and ValueGuided Waypoint Generation (VGWG). HSTE encodes long-term global memory using topological graphs and captures shortterm local details via grid maps. MGAF aligns these dualmap representations with instructions through geometry-aware knowledge fusion. The resulting representations are iteratively refined through pretraining tasks. During reasoning, VGWG generates Guided Attention Heatmaps (GAHs) to explicitly model environment-instruction relevance and optimize waypoint selection. Extensive comparative experiments and performance analyses are conducted, demonstrating that ST-Booster outperforms existing state-of-the-art methods, particularly in complex, disturbance-prone environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。