多智能体强化学习让无人机竞速超越人类,且更安全
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

- 用自对弈训练多无人机协同竞速,学会预判与避障
- 速度超22米/秒时仍比人类冠军快,碰撞率降低50%
- 可零样本适配人类交互,适合机器人协作研究
自主系统在孤立或仿真环境中已实现超人表现,但在共享动态真实空间中仍显脆弱。这源于物理应用中普遍采用的单智能体范式,忽略其他行为体或将其视为环境噪声,难以实现有效协调。本文展示多智能体强化学习为真实世界交互提供了必要的安全框架。以高速四旋翼竞速为高风险测试平台,训练智能体在不同数量竞速者下应对复杂空气动力学相互作用和策略性机动。通过联赛式自对弈,智能体发展出前瞻性行为,包括主动避碰、超车及处理多智能体物理交互(如气流下洗)。其在超过22米/秒的速度下,多玩家比赛中击败冠军级人类飞行员,同时相比最先进单智能体基线碰撞率降低50%。关键的是,使用多样化人工智能体训练使模型具备零样本泛化能力,可安全与人类交互。结果表明,实现机器人稳健共存的路径不在于孤立的安全约束,而在于多智能体互动的严格要求。
原文摘要 · Abstract (English)
Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This failure stems from the dominant single-agent paradigm for physical applications, where other actors are ignored or treated as environmental noise, preventing effective coordination. Here we show that multi-agent reinforcement learning provides the essential safety scaffolding required for real-world interaction. Using high-speed quadrotor racing as a high-stakes testbed, we train agents to navigate complex aerodynamic interactions and strategic maneuvering with a variable number of racers. Through league-based self-play, agents evolve sophisticated anticipatory behaviors, including proactive collision avoidance, overtaking, and handling multi-agent physical interactions, including aerodynamic downwash. Our agents outperform a champion-level human pilot in multi-player races at speeds exceeding 22 m/s, while simultaneously reducing collision rates by 50 % compared to state-of-the-art single-agent baselines. Crucially, training with diverse artificial agents enables zero-shot generalization to safer human interaction. These results suggest that the path to robust robotic co-existence lies not in isolated safety constraints, but in the rigorous demands of multi-agent interaction. Multimedia materials are available at: https://rpg.ifi.uzh.ch/marl
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。