arXiv:2410.08500cs.ROcs.AI2024-10被引 23

用空间语义地图增强大模型在航拍导航中的推理能力

Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation

  • 构建语义-拓扑-度量地图,动态融合视觉与语言信息
  • 零样本下航拍导航成功率提升26.8%(简单任务)
  • 适合无人机导航、多模态智能系统研究者参考

航拍视觉语言导航(Aerial VLN)是一项新任务,使无人机通过自然语言指令和视觉线索在室外环境中导航。由于空中场景中空间关系复杂,该任务仍具挑战性。本文提出一种无需训练的零样本框架,利用大语言模型(LLM)进行动作预测。我们设计了一种新型语义-拓扑-度量表示(STMR),将与指令相关的语义掩码投影到俯视图地图上,动态生成包含周围地标空间与拓扑信息的地图。每一步从增长的俯视图中提取以无人机为中心的局部地图,并转换为带距离度量的矩阵表示,作为文本提示输入给LLM,以响应语言指令进行动作预测。在真实与仿真环境中的实验验证了方法的有效性与鲁棒性,在简单与复杂导航任务上分别相较当前最优方法提升26.8%和5.8%的成功率。数据集与代码即将开源。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation (VLN) is a novel task enabling Unmanned Aerial Vehicles (UAVs) to navigate in outdoor environments through natural language instructions and visual cues. However, it remains challenging due to the complex spatial relationships in aerial scenes.In this paper, we propose a training-free, zero-shot framework for aerial VLN tasks, where the large language model (LLM) is leveraged as the agent for action prediction. Specifically, we develop a novel Semantic-Topo-Metric Representation (STMR) to enhance the spatial reasoning capabilities of LLMs. This is achieved by extracting and projecting instruction-related semantic masks onto a top-down map, which presents spatial and topological information about surrounding landmarks and grows during the navigation process. At each step, a local map centered at the UAV is extracted from the growing top-down map, and transformed into a ma trix representation with distance metrics, serving as the text prompt to LLM for action prediction in response to the given instruction. Experiments conducted in real and simulation environments have proved the effectiveness and robustness of our method, achieving absolute success rate improvements of 26.8% and 5.8% over current state-of-the-art methods on simple and complex navigation tasks, respectively. The dataset and code will be released soon.

航拍导航大模型推理空间表示视觉语言导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。