arXiv:2606.01205cs.RO2026-06被引 3

用想象生成未来画面,让无人机更智能地听懂指令飞行

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

论文配图:ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning
图 1 · 摘自论文原文
  • 通过潜空间视频扩散模型预演环境变化,生成指令对应的未来视觉
  • 在真实飞行中实现92.3%成功率,优于现有方法且仅需1.3亿参数
  • 适合追求高鲁棒性无人机导航的科研与工业应用

无人机视觉语言导航(VLN)需在部分可观测条件下,将自由文本指令转化为六自由度(6-DoF)飞行路径。尽管视觉-语言-动作(VLA)模型具备语义推理能力,但因几何不一致与动力学失配而易出错。为此,我们提出ImagineUAV,一种基于多级世界-动作建模的想象驱动框架。该框架不直接回归动作,而是利用潜空间视频扩散模型生成条件化未来观测,显式模拟环境演化,并通过动作提取器推断6-DoF运动。随后,动力学规划器将估计结果优化为无碰撞轨迹。此外,一步蒸馏推理流程保障实时执行。仅1.3亿参数下,ImagineUAV在基准测试与真实飞行中均超越先前的VLN与VLA基线,验证了想象驱动航法的实用性。

原文摘要 · Abstract (English)

Vision-language navigation (VLN) for UAVs demands grounding free-form instructions into 6-DoF flight under partial observability. While Vision-Language-Action (VLA) models excel at semantic reasoning, they suffer from brittleness due to geometric inconsistency and dynamics mismatch. To address this, we propose ImagineUAV, an imagination-driven framework leveraging cascaded world-action modeling. Instead of direct regression, ImagineUAV employs a latent video diffusion model to generate instruction-conditioned future observations, explicitly imagining environmental evolution, from which 6-DoF motions are inferred via an action extractor. A kinodynamic planner then refines these estimates into collision-free trajectories. Additionally, a step-distilled inference pipeline ensures real-time execution. With only 1.3B parameters, ImagineUAV outperforms prior VLN and VLA baselines on benchmarks and real-world flights, validating the practicality of imagination-driven aerial navigation.

无人机导航视觉语言扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。