arXiv:2603.14363cs.CVcs.AI2026-03被引 13

让无人机通过简单指令自主飞行,无需依赖外部探测器。

AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control

  • 直接从视觉和模糊指令生成连续控制信号,端到端简化流程。
  • 在未见场景中成功率达基线三倍,超越现有方法。
  • 适合追求高自主性无人机导航的研究者与开发者。

无人机视觉语言导航(VLN)需在动态3D环境中进行复杂视觉解析与连续控制。现有分层方法依赖密集的真值引导或辅助目标检测器,造成语义断层并限制真正自主性。本文提出AerialVLA,一种极简端到端的视觉-语言-动作框架,将原始视觉观测与模糊语言指令直接映射为连续物理控制信号。首先,设计轻量双视角感知策略,在减少视觉冗余的同时保留前向导航与精准定位的关键线索,有利于未来仿真到现实的迁移。为恢复真正自主性,引入仅依赖机载传感器的模糊方向提示机制,彻底消除对密集真值引导的依赖。最终,构建统一控制空间,整合三维运动自由度(3-DoF)连续运动指令与内在着陆信号,使智能体无需外部检测器即可实现精准着陆。在TravelUAV基准上的大量实验表明,AerialVLA在已见环境达到最先进性能;在未见场景中,其成功率接近领先基线的三倍,验证了极简、以自主性为中心的范式能捕捉更鲁棒的视觉-运动表征,优于复杂的模块化系统。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) for Unmanned Aerial Vehicles (UAVs) demands complex visual interpretation and continuous control in dynamic 3D environments. Existing hierarchical approaches rely on dense oracle guidance or auxiliary object detectors, creating semantic gaps and limiting genuine autonomy. We propose AerialVLA, a minimalist end-to-end Vision-Language-Action framework mapping raw visual observations and fuzzy linguistic instructions directly to continuous physical control signals. First, we introduce a streamlined dual-view perception strategy that reduces visual redundancy while preserving essential cues for forward navigation and precise grounding, which additionally facilitates future simulation-to-reality transfer. To reclaim genuine autonomy, we deploy a fuzzy directional prompting mechanism derived solely from onboard sensors, completely eliminating the dependency on dense oracle guidance. Ultimately, we formulate a unified control space that integrates continuous 3-Degree-of-Freedom (3-DoF) kinematic commands with an intrinsic landing signal, freeing the agent from external object detectors for precision landing. Extensive experiments on the TravelUAV benchmark demonstrate that AerialVLA achieves state-of-the-art performance in seen environments. Furthermore, it exhibits superior generalization in unseen scenarios by achieving nearly three times the success rate of leading baselines, validating that a minimalist, autonomy-centric paradigm captures more robust visual-motor representations than complex modular systems.

无人机导航端到端视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。