arXiv:2502.01932cs.ROcs.AI2025-02NeurIPS被引 3

多无人机排球测试平台,融合运动控制与策略博弈。

VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play

  • 构建可协作可竞争的多无人机排球系统,支持3D敏捷机动
  • 分层策略在3对3任务中胜率69.5%,优于现有方法
  • 首次实现纯仿真训练策略零样本部署到真实无人机

机器人体育以明确目标、规则清晰和动态交互为特征,是展示具身智能的理想场景。本文提出VolleyBots,一个全新的机器人体育测试平台,允许多个无人机在物理动力学环境下协作与对抗排球比赛。VolleyBots在统一平台上集成三大特性:竞争与合作并存的游戏模式、回合制交互结构以及高敏捷3D机动能力。这些交织特性构成一个复杂问题,同时涉及运动控制与策略决策,且无专家示范可用。我们设计了从单机训练到多机协同与对抗的完整任务体系,并提供了代表性强化学习(RL)、多智能体强化学习(MARL)及博弈论算法的基线评估。仿真结果显示,在单智能体任务中,在线策略(on-policy)RL方法优于离线策略(off-policy)方法,但在融合运动控制与策略决策的复杂任务中,两类方法均表现受限。我们进一步设计分层策略,在3对3任务中击败最强基线,取得69.5%胜率,展现出应对低层控制与高层策略耦合问题的潜力。为验证模拟到现实的可行性,我们成功实现完全在仿真中训练的策略零样本部署至真实无人机上。

原文摘要 · Abstract (English)

Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering. These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations. We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play. We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy. To highlight VolleyBots' sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones.

多无人机强化学习运动控制仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。