用物理仿真构建可解释的羽毛球自对弈环境,让智能体学会联动击球与移动。
ShuttleArena: Interpretable Self-Play in Physics-Based Badminton

- 设计耦合飞行、拦截与回位的物理仿真环境,支持策略解耦分析。
- 自对弈训练使智能体在单回合中实现击球与防守协同优化。
- 结果可解释,适合研究对抗性运动中的战术决策与位置策略。
羽毛球是游戏人工智能的紧凑而具挑战性的领域:选手必须选择符合物理规律的球路,预判对手拦截,并恢复至依赖对手下一步反应的位置。核心难点在于击球与回位不可分离:最佳回位取决于击球引发的对手反应,而击球价值又取决于击球者能否覆盖对方回击。本文提出ShuttleArena,一个基于物理的单人羽毛球自对弈环境,整合连续球飞行、选手拦截、结构化击球生成与击球后回位。策略采用角色条件输出:接球方使用掩码拦截选择,击球方则分解为方位角、仰角、速度与回位目标的因子化动作,实现可解释的战术探查。回合以单次发球-回击为准,训练采用近端策略优化(PPO)自对弈,对抗分阶段检查点对手池,奖励稀疏且基于回合结束结果,同时引入因子特定的回位更新。通过冻结检查点测试、受控战术探查、回位消融实验、定性轨迹演示及人类数据合理性验证,结果显示性能提升显著,且击球几何与回位行为随对手变化呈现可解释性调整。学习策略展现出典型的羽毛球结构特征,同时体现模拟器抽象,回位干预表明学习到的回位行为具有竞争性重要性。结果表明,基于物理的球类运动是交互式数字娱乐人工智能的有力测试平台,因其要求智能体协调执行、站位与对手相对战术价值。
原文摘要 · Abstract (English)
Badminton is a compact but challenging domain for game AI: a player must choose a physically feasible shuttle trajectory, anticipate the opponent's interception, and recover to a court position whose value depends on the opponent's next response. The central challenge is that shot selection and recovery are not separable: the best recovery depends on the shot-induced opponent response, while the value of the shot depends on whether the hitter can cover the reply. This paper presents ShuttleArena, a physics-based singles badminton self-play environment that couples continuous shuttle flight, player interception, structured shot generation, and post-shot recovery. The policy uses role-conditioned outputs: a masked interception choice on receiver turns and a factorized hitter action over shot azimuth, shot elevation, shot speed, and recovery target, enabling interpretable tactical probes. Episodes are single rallies rather than full scored games, and training uses Proximal Policy Optimization (PPO) self-play against a staged checkpoint opponent pool with sparse terminal rally-outcome rewards and a factor-specific recovery update. Evaluation with frozen checkpoint play, controlled tactical probes, recovery ablations, qualitative rollouts, and a human-data sanity check shows competitive improvement together with interpretable opponent-conditioned changes in shot geometry and recovery behavior. The learned policies produce recognizable badminton-like structure while also reflecting the abstractions of the simulator, and the recovery intervention shows that learned recovery behavior is competitively important. These results suggest that physics-based racket sports are a useful testbed for interactive digital entertainment AI because they require agents to coordinate execution, positioning, and opponent-relative tactical value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。