让AI导航更智能:主动感知+聚焦思考,提升视觉语言导航效率
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
- 用大模型协同生成目标视觉查询,主动获取关键信息
- 通过分支多样搜索选出高价值路径点,减少无效推理
- 无需训练即可在多个基准上达到顶尖零样本表现
视觉-语言导航(VLN)要求智能体准确感知复杂视觉环境,并对导航指令与历史信息进行推理。现有方法被动处理冗余视觉输入,对所有历史上下文一视同仁,导致感知低效、推理分散。为此,我们提出无需训练的渐进式框架ProFocus,通过大语言模型(LLMs)与视觉语言模型(VLMs)协作,统一主动感知与聚焦推理。在主动感知方面,ProFocus将全景观测转化为结构化的以我为中心语义地图,使协调智能体识别决策所需缺失视觉信息,并生成带关注区域的目标视觉查询,引导感知智能体获取必要观测。在聚焦推理方面,提出分支多样蒙特卡洛树搜索(BD-MCTS),从大量历史候选点中筛选出前-k个高价值路径点,决策智能体仅聚焦于这些路径点相关的历史上下文,而非平均处理全部历史。大量实验验证了ProFocus的有效性,在R2R和REVERIE基准上实现零样本方法中的最佳性能。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual inputs and treat all historical contexts indiscriminately, resulting in inefficient perception and unfocused reasoning. To address these challenges, we propose \textbf{ProFocus}, a training-free progressive framework that unifies \underline{Pro}active Perception and \underline{Focus}ed Reasoning through collaboration between large language models (LLMs) and vision-language models (VLMs). For proactive perception, ProFocus transforms panoramic observations into structured ego-centric semantic maps, enabling the orchestration agent to identify missing visual information needed for reliable decision-making, and to generate targeted visual queries with corresponding focus regions that guide the perception agent to acquire the required observations. For focused reasoning, we propose Branch-Diverse Monte Carlo Tree Search (BD-MCTS) to identify top-$k$ high-value waypoints from extensive historical candidates. The decision agent focuses reasoning on the historical contexts associated with these waypoints, rather than considering all historical waypoints equally. Extensive experiments validate the effectiveness of ProFocus, achieving state-of-the-art performance among zero-shot methods on R2R and REVERIE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。