arXiv:2509.14082cs.RO2025-09被引 2

用扩散模型生成第一视角视频,让无人机低成本训练更智能。

FlightDiffusion: Revolutionising Autonomous Drone Training with Diffusion Models Generating FPV Video

  • 用扩散模型从单帧生成带动作空间的逼真视频序列。
  • 生成轨迹误差小(位置RMSE 0.28m,朝向RMSE 0.24rad),可直接用于训练。
  • 适合想低成本训练无人机或研究视觉导航的研究者。

我们提出FlightDiffusion,一个基于扩散模型的框架,用于从第一人称视角(FPV)视频训练自主无人机。该模型能从单帧生成包含对应动作空间的逼真视频序列,支持动态环境中的推理式导航。除了直接策略学习外,它还利用生成能力合成多样化的FPV轨迹和状态-动作对,实现大规模训练数据的低成本构建。评估表明,生成轨迹在物理上合理且可执行:位置均方根误差为0.28米,方向均方根误差为0.24弧度。该方法提升了策略学习效果与数据集可扩展性,在下游导航任务中表现更优。仿真环境中表现鲁棒,轨迹更平滑,并能适应未知条件。方差分析显示模拟与真实性能无显著差异(F(1,16)=0.394, p=0.541),成功率分别为M=0.628(SD=0.162)和M=0.617(SD=0.177),证明了强的模拟到现实迁移能力。生成的数据集为未来无人飞行器研究提供了宝贵资源。本工作引入基于扩散模型的推理范式,统一了导航、动作生成与数据合成在空中机器人中的应用。

原文摘要 · Abstract (English)

We present FlightDiffusion, a diffusion-model-based framework for training autonomous drones from first-person view (FPV) video. Our model generates realistic video sequences from a single frame, enriched with corresponding action spaces to enable reasoning-driven navigation in dynamic environments. Beyond direct policy learning, FlightDiffusion leverages its generative capabilities to synthesize diverse FPV trajectories and state-action pairs, facilitating the creation of large-scale training datasets without the high cost of real-world data collection. Our evaluation demonstrates that the generated trajectories are physically plausible and executable, with a mean position error of 0.25 m (RMSE 0.28 m) and a mean orientation error of 0.19 rad (RMSE 0.24 rad). This approach enables improved policy learning and dataset scalability, leading to superior performance in downstream navigation tasks. Results in simulated environments highlight enhanced robustness, smoother trajectory planning, and adaptability to unseen conditions. An ANOVA revealed no statistically significant difference between performance in simulation and reality (F(1, 16) = 0.394, p = 0.541), with success rates of M = 0.628 (SD = 0.162) and M = 0.617 (SD = 0.177), respectively, indicating strong sim-to-real transfer. The generated datasets provide a valuable resource for future UAV research. This work introduces diffusion-based reasoning as a promising paradigm for unifying navigation, action generation, and data synthesis in aerial robotics.

无人机扩散模型数据生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。