用潜在空间博弈让机器人拳击更智能且不摔倒。
RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing

- 把拳击任务转为潜在空间的零和博弈,避免动作崩溃。
- 实测胜率与击打效率显著优于直接在原始动作空间探索。
- 适合研究人形机器人对抗策略或强化学习落地应用者。
在接触频繁、动态剧烈的任务如拳击中,实现人形机器人的高水平竞技智能与身体敏捷性仍是重大挑战。尽管多智能体强化学习为策略交互提供了理论框架,但其直接应用于未结构化的原始运动空间会导致关节级物理崩溃,阻碍有效战斗策略的形成。为解决战略探索与物理可行性之间的根本矛盾,我们提出一种新型双人潜在空间零和马尔可夫博弈。在标准正则性和近似最优响应假设下,证明潜在形式在解码器可达的动作流形上诱导出等价博弈,为自对弈动态提供近似纳什解释。为此,我们构建了分层框架RoboStriker:首先将预定义拳击动作的追踪知识提炼为拓扑受限的潜在流形;再通过潜在空间神经虚构自博弈驱动多智能体协同进化。大量实验表明,该结构化潜在空间内的博弈性能远超原始动作空间探索。通过预训练动作解码器约束战略探索,RoboStriker显著减少原始动作空间方法中的灾难性失衡问题,并在胜负率与击打效率上取得更优表现。最终,我们成功在真实人形机器人上部署并验证了所学战斗策略。代码、视频及补充材料见RoboStriker。
原文摘要 · Abstract (English)
Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challenge, particularly in contact-rich and highly dynamic tasks such as boxing. While Multi-Agent Reinforcement Learning offers a principled framework for strategic interaction, its direct application to unstructured raw motor spaces inevitably leads to joint-level physical collapse, preventing the emergence of any viable combat tactics. To resolve this fundamental conflict between strategic exploration and physical feasibility, we formulate the humanoid combat task as a novel two-player latent-space zero-sum Markov game. Under standard regularity and approximate best-response assumptions, we show that the latent formulation induces an equivalent game over the decoder-reachable action manifold, providing an approximate-Nash interpretation of the resulting self-play dynamics. To instantiate this theoretical formulation, we propose RoboStriker, a hierarchical framework that decouples high-level reasoning from low-level execution. It first distills the tracking expertise of predefined boxing motions into a topologically bounded latent manifold. This structured latent foundation subsequently drives multi-agent co-evolution via Latent-Space Neural Fictitious Self-Play. Extensive experimental results demonstrate that gaming within this structured latent space substantially outperforms direct exploration. By constraining strategic exploration through a pretrained motion decoder, RoboStriker substantially reduces the catastrophic balance failures observed in raw action-space methods and achieves superior tactical performance in both competitive win rates and striking efficiency. Finally, we successfully deploy and validate our learned combat policies on real-world humanoid robots. Our code and video and supplementary materials are available at RoboStriker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。