arXiv:2601.12742cs.ROcs.AI2026-01被引 3

让无人机用自然语言找物,又快又准还省电

AirHunt: Bridging VLM Semantics and Continuous Planning for Efficient Aerial Object Navigation

  • 双路异步架构让视觉语言模型与路径规划协同工作
  • 实测成功率更高,导航误差更小,飞行时间减少30%以上
  • 适合需要快速定位开放集物体的复杂户外场景

大型视觉-语言模型(VLM)提供了丰富的语义理解能力,使无人机可通过自然语言指令在室外环境中搜索开放集物体。然而,以往系统因VLM推理频率与实时规划存在数量级差异,且缺乏三维场景理解能力,难以集成到实际空中系统中。此外,它们缺乏统一机制来平衡语义引导与运动效率。为此,我们提出AirHunt,一个高效定位开放集物体的空中目标导航系统,具备零样本泛化能力。AirHunt采用双路异步架构,实现VLM语义推理与连续路径规划的无缝融合,支持持续飞行与动态语义引导。我们提出主动双任务推理模块,利用几何与语义冗余实现选择性VLM调用;并设计语义-几何一致规划模块,在统一框架内动态协调语义优先级与运动效率,适应环境异质性。我们在多种任务与环境下评估AirHunt,结果表明其成功率更高、导航误差更小、飞行时间更短,优于当前最优方法。真实世界实验进一步验证了其在复杂挑战环境中的实用性。代码与数据集将在发表前公开。

原文摘要 · Abstract (English)

Recent advances in large Vision-Language Models (VLMs) have provided rich semantic understanding that empowers drones to search for open-set objects via natural language instructions. However, prior systems struggle to integrate VLMs into practical aerial systems due to orders-of-magnitude frequency mismatch between VLM inference and real-time planning, as well as VLMs' limited 3D scene understanding. They also lack a unified mechanism to balance semantic guidance with motion efficiency in large-scale environments. To address these challenges, we present AirHunt, an aerial object navigation system that efficiently locates open-set objects with zero-shot generalization in outdoor environments by seamlessly fusing VLM semantic reasoning with continuous path planning. AirHunt features a dual-pathway asynchronous architecture that establishes a synergistic interface between VLM reasoning and path planning, enabling continuous flight with adaptive semantic guidance that evolves through motion. Moreover, we propose an active dual-task reasoning module that exploits geometric and semantic redundancy to enable selective VLM querying, and a semantic-geometric coherent planning module that dynamically reconciles semantic priorities and motion efficiency in a unified framework, enabling seamless adaptation to environmental heterogeneity. We evaluate AirHunt across diverse object navigation tasks and environments, demonstrating a higher success rate with lower navigation error and reduced flight time compared to state-of-the-art methods. Real-world experiments further validate AirHunt's practical capability in complex and challenging environments. Code and dataset will be made publicly available before publication.

无人机导航视觉语言模型路径规划语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。