通过分层语义空间表示,提升视觉语言导航中智能体对环境的理解能力。
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
- 分层架构融合文本语义与深度空间感知,多尺度理解环境
- 在REVERIE、R2R等离散任务上超越基线,连续任务上泛化更强
- 适合研究视觉语言导航、多模态智能体的开发者参考
从自然语言指令中导航未知环境仍是视点导向视觉语言导航(VLN)中智能体面临的挑战。人类在室内导航时会自然地将具体语义知识嵌入空间布局。尽管已有工作引入多种环境表征以提升推理能力,但辅助模态常被简单拼接至RGB特征,未能充分发挥各模态的独立贡献。本文提出分层语义理解与空间感知(SUSA)架构,使智能体能在多个尺度上感知并定位环境。具体而言,文本语义理解(TSU)模块生成视图级描述,支持局部动作预测,捕捉细粒度语义并缩小指令与环境间的模态差异;互补地,深度增强空间感知(DSP)模块逐步构建轨迹级深度探索地图,提供全局空间布局的粗粒度表征。大量实验表明,SUSA的分层表征增强显著提升了基准模型在离散VLN基准(REVERIE、R2R、SOON)上的导航性能,并在连续版本R2R-CE上展现出更强泛化能力。
原文摘要 · Abstract (English)
Navigating unseen environments from natural language instructions remains challenging for egocentric agents in Vision-and-Language Navigation (VLN). Humans naturally ground concrete semantic knowledge within spatial layouts during indoor navigation. Although prior work has introduced diverse environment representations to improve reasoning, auxiliary modalities are often naively concatenated with RGB features, which underutilizes each modality's distinct contribution. We propose a hierarchical Semantic Understanding and Spatial Awareness (SUSA) architecture to enable agents to perceive and ground environments at multiple scales. Specifically, the Textual Semantic Understanding (TSU) module supports local action prediction by generating view-level descriptions, capturing fine-grained semantics and narrowing the modality gap between instructions and environments. Complementarily, the Depth Enhanced Spatial Perception (DSP) module incrementally builds a trajectory-level depth exploration map, providing a coarse-grained representation of global spatial layout. Extensive experiments show that the hierarchical representation enrichment of SUSA significantly improves navigation performance over the baseline on discrete VLN benchmarks (REVERIE, R2R, and SOON) and generalizes better to the continuous R2R-CE benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。