arXiv:2606.06147cs.AI2026-06被引 1

用世界模型预判未来,让无人机在复杂城市中更聪明地飞行

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

论文配图:WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation
图 1 · 摘自论文原文
  • 基于世界模型构建双分支生成框架,同步预测未来画面与飞行动作
  • 在复杂城市峡谷场景中,新环境下的导航成功率提升显著
  • 适合研究具身智能、视觉语言导航与无人机自主决策的学者

端到端视觉-语言-动作(VLA)模型在无人机导航中展现出潜力。然而,现有方法通常依赖历史观测直接预测动作,在密集城市环境中因严重遮挡和急转弯导致视角剧烈变化时表现不佳。我们认为,具备‘想象’未来状态的能力——即世界模型的核心能力——对部分可观测条件下的鲁棒决策至关重要。为此,我们构建了一个名为Urban Canyon Traversal Benchmark的挑战性基准,专门评估在严重遮挡和剧烈视角变换场景下的空间理解能力。在此基础上,我们提出WorldFly,一种基于世界模型的新型VLA框架,采用双分支耦合流匹配机制,联合生成未来视频预测与导航动作,从而通过空间想象显式引导智能体策略。在该基准上的大量评估表明,WorldFly在未见环境中显著优于其他基线模型,验证了将世界模型融入具身空中代理的有效性。

原文摘要 · Abstract (English)

End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict actions, often struggling in dense urban environments where severe occlusions and sharp turns result in drastic viewpoint transitions. We argue that the ability to "imagine" future states -- inherent in World Models -- is critical for robust decision-making under such partial observability. To address this, we construct a challenging Urban Canyon Traversal Benchmark, specifically designed to evaluate spatial understanding in scenarios characterized by severe occlusions and drastic viewpoint transitions. To this end, we propose WorldFly, a novel world-model-based VLA framework that employs a dual-branch coupled flow matching mechanism to jointly generate future video predictions and navigation actions, thereby explicitly guiding the agent's policy via spatial imagination. Extensive evaluations on our benchmark demonstrate that WorldFly outperforms other baselines, particularly in unseen environments, validating the effectiveness of integrating world models into embodied aerial agents.

无人机导航世界模型视觉语言动作具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。