arXiv:2604.21355cs.RO2026-04

让机器人打斗动作切换更流畅稳定,避免因状态突变导致摔倒。

RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting

论文配图:RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting
图 1 · 摘自论文原文
  • 用统一策略融合多种打斗技能,通过随机化训练提升鲁棒性。
  • 实现在仿真与真实机器人(Unitree G1)上长时间连续打斗,支持任意中断或切换。
  • 适合需要复杂动作衔接的仿人机器人应用,如竞技对抗或智能训练。

仿人机器人在多种任务中已展现出出色的运动能力,但实现类人长期动态打斗仍面临巨大挑战,主要源于对敏捷性与稳定性的严苛要求。尽管模仿学习可使机器人执行类人打斗技能,现有方法多依赖多个单技能策略切换或使用通用策略模仿参考动作,导致技能间过渡时因初始与终态不匹配产生域外扰动,引发行为不连贯或不稳定。本文提出RPG(Robust Policy Gating),一种混合专家策略框架,用于实现仿人机器人多技能间的平滑稳定过渡。该方法引入运动过渡随机化与时间随机化,训练统一策略以生成兼具敏捷性、稳定性与过渡平滑性的打斗动作。此外,设计了集成行走/奔跑与打斗技能的控制流水线,支持任意时长的类人持续战斗,并可在任意时刻无缝中断或切换动作策略。大量仿真实验验证了框架有效性,真实世界部署于Unitree G1机器人进一步证明其鲁棒性与实用性。

原文摘要 · Abstract (English)

Humanoid robots have demonstrated impressive motor skills in a wide range of tasks, yet whole-body control for humanlike long-time, dynamic fighting remains particularly challenging due to the stringent requirements on agility and stability. While imitation learning enables robots to execute human-like fighting skills, existing approaches often rely on switching among multiple single-skill policies or employing a general policy to imitate input reference motions. These strategies suffer from instability when transitioning between skills, as the mismatch of initial and terminal states across skills or reference motions introduces out-of-domain disturbances, resulting in unsmooth or unstable behaviors. In this work, we propose RPG, a hybrid expert policy framework, for smooth and stable humanoid multi-skills transition. Our approach incorporates motion transition randomization and temporal randomization to train a unified policy that generates agile fighting actions with stability and smoothness during skill transitions. Furthermore, we design a control pipeline that integrates walking/running locomotion with fighting skills, allowing humanlike long-time combat of arbitrary duration that can be seamlessly interrupted or transit action policies at any time. Extensive experiments in simulation demonstrate the effectiveness of the proposed framework, and real-world deployment on the Unitree G1 humanoid robot further validates its robustness and applicability.

仿人机器人动作过渡模仿学习多技能切换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。