让无人机通过自然语言指令完成复杂任务,提升空中导航的准确性与效率。
SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning
- 将3D几何信息融入视觉输入,增强大模型对空间的理解能力。
- 在2.5D场景中成功率提升25.7%,3D复杂场景中提升39.3%。
- 适合需要精准空域导航的工业巡检、灾害救援等实际应用。
自然语言驱动的自主导航是实现具身智能的关键一步,可使无人机在工厂、家庭等环境中执行复杂任务。然而,当前零样本视觉-语言模型(VLM)缺乏精确的空间推理能力,常生成模糊或不可行的指令,且现有视觉-语言导航方法主要针对地面机器人,难以适应小规模、杂乱环境中的三维空中任务。本文提出SoraNav框架,实现无人机任务导向的零样本VLM推理。为弥合空间语义鸿沟,引入多模态视觉标注(MVA),将3D几何先验编码至VLM的2D视觉输入;为避免错误指令导致死胡同和重复探索,设计自适应决策策略(ADM),结合探索历史验证并动态切换至基于几何的探索方式。部署于定制的PX4微无人机平台,实验表明该方法显著优于现有基线:在2.5D场景中成功率达25.7%提升,导航效率(SPL)提升17.3%;在复杂3D场景中,成功率达到39.3%提升,SPL提升24.7%。
原文摘要 · Abstract (English)
Autonomous navigation under natural language instructions represents a crucial step toward embodied intelligence, enabling complex task execution in environments ranging from industrial facilities to domestic spaces. However, language-driven 3D navigation for Unmanned Aerial Vehicles (UAVs) requires precise spatial reasoning, a capability inherently lacking in current zero-shot Vision-Language Models (VLMs) which often generate ambiguous outputs and cannot guarantee geometric feasibility. Furthermore, existing Vision-Language Navigation (VLN) methods are predominantly tailored for 2.5D ground robots, rendering them unable to generalize to the unconstrained 3D spatial reasoning required for aerial tasks in small-scale, cluttered environments. In this paper, we present SoraNav, a novel framework enabling zero-shot VLM reasoning for UAV task-centric navigation. To address the spatial-semantic gap, we introduce Multi-modal Visual Annotation (MVA), which encodes 3D geometric priors directly into the VLM's 2D visual input. To mitigate hallucinated or infeasible commands, we propose an Adaptive Decision Making (ADM) strategy that validates VLM proposals against exploration history, seamlessly switching to geometry-based exploration to avoid dead-ends and redundant revisits. Deployed on a custom PX4-based micro-UAV, SoraNav demonstrates robust real-world performance. Quantitative results show our approach significantly outperforms state-of-the-art baselines, increasing Success Rate (SR) by 25.7% and navigation efficiency (SPL) by 17.3% in 2.5D scenarios, and achieving improvements of 39.3% (SR) and 24.7% (SPL) in complex 3D scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。