用渐进式训练让多无人机竞速更高效、更智能。
Curriculum-Based Iterative Self-Play for Scalable Multi-Drone Racing
- 设计渐进难度课程+自对弈机制,提升训练效率。
- 竞速平均速度接近现有方法两倍,成功率高。
- 适合研究多智能体协作与高速自主系统的人看。
在高速竞争环境中协调多个自主代理是重大工程挑战。本文提出CRUISE(基于课程的迭代自对弈多无人机竞速框架),通过将渐进式难度课程与高效自对弈机制相结合,突破了传统方法的可扩展性瓶颈,有效促进鲁棒竞争行为的生成。在高保真仿真中,基于真实四旋翼动力学验证,所获策略显著优于标准强化学习基线和先进博弈论规划器。CRUISE实现近两倍于规划器的平均竞速速度,保持高成功率,并在智能体密度增加时仍具强可扩展性。消融实验表明,课程结构是性能跃升的关键因素。该方法为动态竞争任务中的自主系统开发提供了可扩展且高效的训练范式,有望推动未来真实场景部署。
原文摘要 · Abstract (English)
The coordination of multiple autonomous agents in high-speed, competitive environments represents a significant engineering challenge. This paper presents CRUISE (Curriculum-Based Iterative Self-Play for Scalable Multi-Drone Racing), a reinforcement learning framework designed to solve this challenge in the demanding domain of multi-drone racing. CRUISE overcomes key scalability limitations by synergistically combining a progressive difficulty curriculum with an efficient self-play mechanism to foster robust competitive behaviors. Validated in high-fidelity simulation with realistic quadrotor dynamics, the resulting policies significantly outperform both a standard reinforcement learning baseline and a state-of-the-art game-theoretic planner. CRUISE achieves nearly double the planner's mean racing speed, maintains high success rates, and demonstrates robust scalability as agent density increases. Ablation studies confirm that the curriculum structure is the critical component for this performance leap. By providing a scalable and effective training methodology, CRUISE advances the development of autonomous systems for dynamic, competitive tasks and serves as a blueprint for future real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。