arXiv:2604.02318cs.ROcs.CV2026-04被引 2

用元认知思维让视觉语言导航更高效,减少重复探索

Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning

  • 构建动态语义地图+历史感知规划,避免反复走回头路
  • 检测探索停滞并用大模型生成纠错规则,效率提升20.7%
  • 适合需要低资源、高鲁棒性的智能体导航场景

无需训练的视觉-语言导航(VLN)代理依赖基础模型理解指令并探索3D环境。但现有方法采用贪心前沿选择和被动空间记忆,导致局部振荡和重复访问等低效行为。我们指出其根源在于缺乏元认知能力:无法监测探索进度、诊断策略失败或自我调整。为此提出MetaNav,集成空间记忆、历史感知规划与反思修正机制。空间记忆构建持久的3D语义地图;历史感知规划惩罚重复访问以提升效率;反思修正检测探索停滞,并调用LLM生成纠正规则,指导未来前沿选择。在GOAT-Bench、HM3D-OVON和A-EQA上的实验表明,MetaNav达到当前最优性能,同时减少20.7%的VLM查询,证明元认知推理显著提升导航鲁棒性与效率。

原文摘要 · Abstract (English)

Training-free Vision-Language Navigation (VLN) agents powered by foundation models can follow instructions and explore 3D environments. However, existing approaches rely on greedy frontier selection and passive spatial memory, leading to inefficient behaviors such as local oscillation and redundant revisiting. We argue that this stems from a lack of metacognitive capabilities: the agent cannot monitor its exploration progress, diagnose strategy failures, or adapt accordingly. To address this, we propose MetaNav, a metacognitive navigation agent integrating spatial memory, history-aware planning, and reflective correction. Spatial memory builds a persistent 3D semantic map. History-aware planning penalizes revisiting to improve efficiency. Reflective correction detects stagnation and uses an LLM to generate corrective rules that guide future frontier selection. Experiments on GOAT-Bench, HM3D-OVON, and A-EQA show that MetaNav achieves state-of-the-art performance while reducing VLM queries by 20.7%, demonstrating that metacognitive reasoning significantly improves robustness and efficiency.

视觉导航元认知大模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。