arXiv:2505.24396cs.RO2025-05被引 8

用强化学习让无人机自主完成高难度特技飞行,实现端到端控制。

Reactive Aerobatic Flight via Reinforcement Learning

  • 直接将飞行状态与动作意图映射为控制指令,端到端优化飞行策略。
  • 首次实现无人机在移动门中连续倒飞并实时避障,展现极强敏捷性。
  • 采用自动课程学习和领域随机化,确保仿真到现实的零样本迁移。

四旋翼无人机虽具高度灵活性,但因固有欠驱动特性及高动态机动的复杂性,其特技飞行潜力尚未完全释放。传统方法将轨迹规划与跟踪控制分离,导致跟踪误差大、计算延迟高且对初始条件敏感,难以应对高敏捷场景。受数据驱动方法启发,我们提出一种基于强化学习的框架,直接从无人机状态和特技意图生成控制指令,消除模块分离,实现端到端策略优化,完成极限特技飞行。为提升训练效率与稳定性,引入自动化课程学习策略,动态调整任务难度。结合领域随机化,实现鲁棒的零样本仿真到现实迁移。在真实世界实验中成功验证,首次展示无人机自主完成连续倒飞并实时穿越移动门,展现出前所未有的敏捷性能。

原文摘要 · Abstract (English)

Quadrotors have demonstrated remarkable versatility, yet their full aerobatic potential remains largely untapped due to inherent underactuation and the complexity of aggressive maneuvers. Traditional approaches, separating trajectory optimization and tracking control, suffer from tracking inaccuracies, computational latency, and sensitivity to initial conditions, limiting their effectiveness in dynamic, high-agility scenarios. Inspired by recent breakthroughs in data-driven methods, we propose a reinforcement learning-based framework that directly maps drone states and aerobatic intentions to control commands, eliminating modular separation to enable quadrotors to perform end-to-end policy optimization for extreme aerobatic maneuvers. To ensure efficient and stable training, we introduce an automated curriculum learning strategy that dynamically adjusts aerobatic task difficulty. Enabled by domain randomization for robust zero-shot sim-to-real transfer, our approach is validated in demanding real-world experiments, including the first demonstration of a drone autonomously performing continuous inverted flight while reactively navigating a moving gate, showcasing unprecedented agility.

无人机强化学习特技飞行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。