让导航模型具备自我意识,实时理解自身位置和任务进度。
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

- 引入结构化推理模块,实现空间与任务双重自知
- 在Habitat上超越现有方法,成功率显著提升
- 无需额外传感器,适合大规模视觉语言训练
视觉-语言导航(VLN)要求智能体将语言指令映射到自身的视觉环境移动中。尽管最先进方法利用视觉语言模型(VLM)进行端到端动作预测,但往往缺乏对智能体、指令与场景关系的显式且可解释的理解。相反,构建场景地图用于启发式规划虽直观,却依赖额外3D传感器,阻碍大规模视觉语言预训练。为此,我们提出AwareVLN,一种赋予导航模型自知推理机制的新框架,使其以全端到端、数据驱动的方式理解自身状态与任务进展。该方法包含两项关键创新:(1) 结构化推理模块,促进空间与任务导向的自知;(2) 带进度划分的自动数据引擎,支持高效训练。在Habitat模拟器多个数据集上的大量实验表明,AwareVLN显著优于以往最先进视觉语言导航方法。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for end-to-end action prediction, they often lack an explicit and explainable understanding of the relationships between the agent, the instruction, and the scene. Conversely, explicitly building a scene map for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large-scale vision-language pre-training. To bridge this gap, we propose AwareVLN, a novel framework that equips the navigation model with a self-aware reasoning mechanism, enabling it to understand the agent's state and task progress in a fully end-to-end and data-driven manner. Our approach features two key innovations: (1) a structural reasoning module that fosters spatial and task-oriented self-awareness, and (2) an automatic data engine with progress division for effective training. Extensive experiments on various datasets in Habitat simulator show our AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods. Project page: https://gwxuan.github.io/AwareVLN/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。