arXiv:2606.28397cs.CVcs.AI2026-06

通过闭环验证提升无人机视觉语言导航的准确性。

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

论文配图:CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation
图 1 · 摘自论文原文
  • 采用闭环推理,执行前验证并修正动作
  • 在城市导航基准上达到32.01%成功率达标率
  • 适合需要高可靠性的无人机导航场景

视觉语言导航(VLN)近年来借助大语言和多模态模型取得进展,使智能体可在未见环境中遵循自然语言指令,无需训练特定导航策略。然而,多数现有方法仍采用开环决策-执行模式,生成的动作很少在执行前被验证或修正,导致空中视觉语言导航中微小动作误差快速累积,引发轨迹偏移与目标丢失。为此,本文提出闭环自验证检索增强推理(CLOSER),一种无训练策略的方法,通过在执行前依次完成动作推理、可靠性验证、定向检索与动作修正,实现闭环操作。我们将其应用于空中视觉语言导航任务,构建CLOSER-VLN框架,包含三个组件:基于信息生成候选动作的分层推理器、评估动作可靠性的多维动作验证器,以及在验证失败时从记忆库中检索目标示例的验证触发式多模态检索器。在CityNav基准测试中,CLOSER-VLN在测试未见划分上达到32.01%的路径成功率(SR)和21.28%的标准化路径长度(SPL),验证了闭环推理的有效性。

原文摘要 · Abstract (English)

Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy. However, most existing VLN methods relying on large models still adopt an open-loop decision-execution approach, where candidate actions are generated from instructions and observations but are rarely verified or corrected before execution. This causes critical issues in aerial VLN, where minor errors in intermediate actions may quickly accumulate into large trajectory deviations and lead to target loss. To address this issue, we propose Closed-loop Self-verified Retrieval-augmented Reasoning (CLOSER), a training-policy-free method that sequentially performs action reasoning, reliability verification, targeted retrieval, and action correction in a closed-loop manner before executing concrete actions. We instantiate the CLOSER in aerial VLN tasks and develop a CLOSER-VLN framework, which is composed of three components: a hierarchical reasoner for generating candidate actions based on available information, a multidimensional action verifier for assessing the reliability of actions generated by the reasoner, and a verification-triggered multimodal retriever for retrieving targeted exemplars from a memory bank only when verification fails. We conduct experimental evaluations on the CityNav benchmark, where CLOSER-VLN achieves 32.01% SR and 21.28% SPL on the test-unseen split, confirming the effectiveness of closed-loop reasoning.

视觉导航闭环推理无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。