arXiv:2602.09657cs.RO2026-02被引 17

AutoFly让无人机在未知户外环境自主导航,靠视觉语言和深度信息决策。

AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

  • 用伪深度编码器从图像中提取空间特征,增强导航推理能力。
  • 在模拟与真实环境中比顶尖模型高3.9%成功率,稳定表现。
  • 新数据集强调自主避障与规划,更适合真实世界无人机应用。

视觉语言导航(VLN)要求智能体结合语言指令与视觉观察来导航环境,是具身智能的核心任务。当前无人机(UAV)的VLN研究依赖详细预设指令引导飞行路径,但在真实户外探索中,环境未知且缺乏详细指引,仅能提供粗粒度位置或方向提示,要求无人机自主进行连续规划与障碍规避。为此,我们提出AutoFly,一种端到端的视觉-语言-动作(VLA)模型,实现无人机自主导航。AutoFly引入伪深度编码器,从RGB输入中生成深度感知特征,提升空间推理能力,并采用渐进式两阶段训练策略,有效对齐视觉、深度与语言表征与动作策略。此外,现有VLN数据集存在根本性缺陷:过度依赖指令执行,缺乏自主决策数据,且真实世界数据不足。为此,我们构建了新的自主导航数据集,通过(1)强调连续避障、自主规划与识别流程的轨迹采集;(2)全面整合真实世界数据,推动范式从指令遵循转向自主行为建模。实验表明,AutoFly在模拟与真实环境中均比现有最先进VLA基线高出3.9%的成功率,表现一致可靠。

原文摘要 · Abstract (English)

Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned aerial vehicles (UAVs) relies on detailed, pre-specified instructions to guide the UAV along predetermined routes. However, real-world outdoor exploration typically occurs in unknown environments where detailed navigation instructions are unavailable. Instead, only coarse-grained positional or directional guidance can be provided, requiring UAVs to autonomously navigate through continuous planning and obstacle avoidance. To bridge this gap, we propose AutoFly, an end-to-end Vision-Language-Action (VLA) model for autonomous UAV navigation. AutoFly incorporates a pseudo-depth encoder that derives depth-aware features from RGB inputs to enhance spatial reasoning, coupled with a progressive two-stage training strategy that effectively aligns visual, depth, and linguistic representations with action policies. Moreover, existing VLN datasets have fundamental limitations for real-world autonomous navigation, stemming from their heavy reliance on explicit instruction-following over autonomous decision-making and insufficient real-world data. To address these issues, we construct a novel autonomous navigation dataset that shifts the paradigm from instruction-following to autonomous behavior modeling through: (1) trajectory collection emphasizing continuous obstacle avoidance, autonomous planning, and recognition workflows; (2) comprehensive real-world data integration. Experimental results demonstrate that AutoFly achieves a 3.9% higher success rate compared to state-of-the-art VLA baselines, with consistent performance across simulated and real environments.

无人机导航视觉语言自主决策具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。