提出新框架实现无人机竞速零样本泛化,兼顾高速性能与鲁棒性。
Bridging Performance and Generalization in Reinforcement Learning for Agile Flight

- 结合学习进度的任务切换与物理启发的轨迹生成器
- 实测在未知赛道上泛化能力提升7.4倍,速度接近顶尖水平
- 适合追求真实场景泛化的强化学习研究者
自主无人机竞速对飞行机器人构成根本性挑战,需在持续执行饱和条件下实现时间最优控制。尽管强化学习已在该领域达到人类水平表现,但现有方法泛化能力差:在特定环境中训练的策略在未见配置下常立即坠毁。这一失败源于敏捷飞行中零样本泛化的内在困难,来自高维任务变化以及高速下安全与性能的强耦合。现有提升泛化的方法严重牺牲飞行速度:策略必须显著降速才能获得有限泛化。本文提出一种基于强化学习的无人机竞速零样本泛化框架。通过任务感知的切换机制与物理启发的程序化赛道生成器,该框架无需测试时适应即可生成快速且鲁棒的通用策略。在真实世界多种未见赛道上验证,本方法相较当前最优方法泛化能力提升7.4倍,同时保持具有竞争力的竞速速度。我们在仿真和真实场景中均验证了结果,包括一个无显式状态估计的视觉驱动端到端控制设置,此前所有方法在此设置下均无法泛化。
原文摘要 · Abstract (English)
Autonomous drone racing is a fundamentally challenging regime for autonomous aerial robots, requiring time-optimal control while operating under persistent actuation saturation. While reinforcement learning (RL) has achieved human-level performance in this domain, current methods fail to generalize; policies trained on specific environments often crash immediately in unseen configurations. This failure reflects the intrinsic difficulty of zero-shot generalization in agile flight, arising from high-dimensional task variation and the tight coupling between safety and performance at high speeds. Existing approaches that improve generalization impose a substantial cost on flight speed: control policies must significantly degrade performance to achieve even modest levels of generalization. In this work, we propose a framework for zero-shot generalization in agile flight for RL-based drone racing. By combining task-aware switching based on learning progress with a physically informed procedural track generator, the framework produces a fast and robust generalist policy without test-time adaptation. Our method achieves strong zero-shot performance across a wide range of unseen racetracks in the real world, demonstrating a 7.4x improvement in generalization over the state-of-the-art approaches, while maintaining competitive racing speeds. We validate our method's results in both simulation and real-world settings, including a challenging vision-based, end-to-end control setting that operates without explicit state estimation, where all prior approaches fail to generalize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。