让无人机像飞行员一样思考,实现长程精细导航。
Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

- 分层架构分离任务推理与实时控制,支持长时间复杂指令
- 在双数据集上实现90%以上多阶段任务完成率和高精度轨迹控制
- 适合需要实时决策的无人机自主导航研究者
语言引导的无人机代理需在执行长程语义指令的同时生成平滑、物理可行的连续飞行命令。现有视觉语言导航(VLN)基准通常采用离散或粗粒度动作,而现有无人机视觉-语言-动作(VLA)任务则聚焦于短时原子操作。为填补这一空白,我们提出 extbf{FLIGHT}:一个细粒度长程指令引导的混合无人机导航与推理基准,包含多阶段指令与密集6-自由度轨迹标注,涵盖两个数据集划分:细粒度VLN与长程流程。为赋予无人机代理实时飞行状态推理与任务规划能力,并兼顾高频精确控制,我们进一步提出 extbf{FLIGHT VLA},一种异步架构:低频流式飞行员视觉语言模型(VLM)负责任务状态推理,高频扩散动作模型负责连续控制,由显式的 extbf{飞行员推理} 文本监督,总结当前飞行状态并预判下一子目标。闭环评估显示,FLIGHT VLA 在我们的 FLIGHT 基准上持续优于代表性 VLN 与 VLA 基线,实现更强的多阶段完成率、子目标遵循度与终端控制表现。训练后的流式飞行员推理VLM还提升了无人机视频理解能力,验证了设计有效性。
原文摘要 · Abstract (English)
Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers. To address this gap in UAV task settings, we introduce \textbf{FLIGHT}, a \textbf{F}ine-grained \textbf{L}ong-horizon \textbf{I}nstruction-\textbf{G}uided benchmark for \textbf{H}ybrid UAV navigation and reasoning \textbf{T}asks, which combines multi-stage instructions with dense 6-DoF trajectory annotations across two dataset splits: Fine-grained VLN and Long-horizon Flow. To endow the UAV agent with the capability of real-time in-flight reasoning over task execution status and mission planning, while simultaneously accommodating high-frequency, real-time precise control, we further propose \textbf{FLIGHT VLA}, an asynchronous architecture that decouples a low-frequency Streaming Pilot Vision-Language Model (VLM) for task-state reasoning from a high-frequency diffusion action model for continuous control, supervised by explicit \textbf{Pilot Reasoning} texts that summarize the current flight state and anticipate the next subgoal. In closed-loop evaluation, FLIGHT VLA consistently surpasses representative VLN and VLA baselines on our FLIGHT benchmarks, achieving stronger multi-stage completion, subgoal adherence, and terminal control. Its trained Streaming Pilot Reasoning VLM further improves UAV video reasoning, validating the effectiveness of our design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。