让指令随环境动态变化,提升智能体导航理解能力
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation

- 将指令视为随环境变化的动态状态,逐步更新语义
- 在REVERIE数据集上未见场景下提升2.68%的SPL指标
- 适合需要精准理解语言与视觉上下文的智能体导航研究
视觉语言导航要求智能体在不断变化的视觉环境中遵循自然语言指令。核心挑战在于语言与观测之间的动态耦合:随着智能体视角和空间上下文的变化,指令含义也随之改变。然而,现有模型通常将指令编码为静态全局表示,难以适应当前视觉上下文。为此,本文将指令理解建模为「指令即状态」变量——一个与决策相关、在词粒度上随感知状态逐步演化的指令状态。感知状态指每一步的观测基础导航上下文。为实现该思想,提出粗到细框架S-EGIU,实现状态条件下的段落激活与词粒度语义细化。粗粒度阶段激活与当前观测语义对齐的指令片段;细粒度阶段通过观测引导的词定位与上下文建模,精炼激活片段内部语义。两者协同使指令状态在导航过程中持续依据感知状态更新。S-EGIU在多个关键指标上表现优异,在REVERIE测试未见场景中取得+2.68% SPL提升,并在多个VLN基准上展现一致效率优势,验证了动态指令-感知耦合的价值。
原文摘要 · Abstract (English)
Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dynamic entanglement between language and observations: the meaning of instruction shifts as the agent's field of view and spatial context evolve. However, many existing models encode the instruction as a static global representation, limiting their ability to adapt instruction meaning to the current visual context. We therefore model instruction understanding as an Instruction-as-State variable: a decision-relevant, token-level instruction state that evolves step by step conditioned on the agent's perceptual state, where the perceptual state denotes the observation-grounded navigation context at each step. To realize this principle, we introduce State-Entangled Environment-Guided Instruction Understanding (S-EGIU), a coarse-to-fine framework for state-conditioned segment activation and token-level semantic refinement. At the coarse level, S-EGIU activates the instruction segment whose semantics align with the current observation. At the fine level, it refines the activated segment through observation-guided token grounding and contextual modeling, sharpening its internal semantics under the current observation. Together, these stages maintain an instruction state that is continuously updated according to the agent's perceptual state during navigation. S-EGIU delivers strong performance on several key metrics, including a +2.68% SPL gain on REVERIE Test Unseen, and demonstrates consistent efficiency gains across multiple VLN benchmarks, underscoring the value of dynamic instruction--perception entanglement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。