arXiv:2512.21887cs.ROcs.AI2025-12被引 9

让无人机通过预测未来视觉生成导航,提升长距离飞行成功率。

Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space

  • 基于历史帧与动作预测未来视角,融合语义与几何先验。
  • 在大规模环境中导航成功率达87.3%,显著优于现有模型。
  • 适合研究无人机自主导航与三维世界建模的学者使用。

无人飞行器(UAV)已成为强大的具身智能体。其核心能力之一是在大规模三维环境中实现自主导航。然而,现有导航策略通常仅优化障碍物规避和轨迹平滑等低层目标,缺乏将高层语义融入规划的能力。为此,我们提出ANWM——一种空中导航世界模型,可基于历史帧与动作预测未来视觉观测,使智能体能够根据语义合理性和导航实用性对候选轨迹进行排序。ANWM在4自由度的无人机轨迹数据上训练,并引入物理启发式模块:未来帧投影(FFP),将历史帧投影至未来视角以提供粗略几何先验。该模块有效缓解了远距离视觉生成中的表征不确定性,捕捉了三维轨迹与自身视角观测之间的映射关系。实验证明,ANWM在长距离视觉预测任务中显著优于现有世界模型,并在大规模环境中将无人机导航成功率提升至87.3%。

原文摘要 · Abstract (English)

Unmanned aerial vehicles (UAVs) have emerged as powerful embodied agents. One of the core abilities is autonomous navigation in large-scale three-dimensional environments. Existing navigation policies, however, are typically optimized for low-level objectives such as obstacle avoidance and trajectory smoothness, lacking the ability to incorporate high-level semantics into planning. To bridge this gap, we propose ANWM, an aerial navigation world model that predicts future visual observations conditioned on past frames and actions, thereby enabling agents to rank candidate trajectories by their semantic plausibility and navigational utility. ANWM is trained on 4-DoF UAV trajectories and introduces a physics-inspired module: Future Frame Projection (FFP), which projects past frames into future viewpoints to provide coarse geometric priors. This module mitigates representational uncertainty in long-distance visual generation and captures the mapping between 3D trajectories and egocentric observations. Empirical results demonstrate that ANWM significantly outperforms existing world models in long-distance visual forecasting and improves UAV navigation success rates in large-scale environments.

无人机视觉预测三维导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。