arXiv:2606.08064cs.RO2026-06

多智能体强化学习让机器人协同跳长绳,能自适应不同跳跃节奏。

Cooperative Long Rope Skipping via Multi-Agent Reinforcement Learning

论文配图:Cooperative Long Rope Skipping via Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 分层强化学习:底层学各自甩绳,上层统一调度协调
  • 在仿真和真实机器人上表现更稳定,跳绳成功率超基线
  • 能应对不同人类玩家的跳绳风格,适合协作运动研究

人类展现出卓越的运动敏捷性,可完成跑步、跳跃等多种动态技能,凸显人形机器人在体育运动中的潜力。长绳跳绳需两名甩绳者协同配合,同时适应不同节奏的跳跃者,是极具挑战性的多智能体协作任务。尽管现有方法在单智能体或无交互场景(如跑步、舞蹈、特技)中取得成功,但涉及多参与者精确协同的任务仍鲜有探索。为此,我们提出Marope框架,一种基于多智能体强化学习(MARL)的人形机器人协同跳长绳方法。该框架采用分层强化学习架构:底层通过多智能体强化学习训练去中心化的甩绳策略;上层训练集中式调度策略以协调底层策略执行。为提升对不同玩家行为风格的泛化能力,框架引入多样化的跳跃策略参与协作训练。我们在Unitree G1人形机器人上于仿真与真实环境进行评估,结果表明,Marope优于多种基线方法,在绳子操控效率与稳定性,以及对多样化玩家的鲁棒性和适应性方面均有显著提升。

原文摘要 · Abstract (English)

Humans exhibit remarkable motor agility, enabling a wide range of dynamic skills such as running and jumping, which highlights the great potential of humanoid robots for athletic locomotion. Among athletic sports, long rope skipping requires two rope turners to cooperatively swing the rope while adapting to a player under different jumping rhythms, making it a meaningful yet challenging task for humanoid robots. Although existing methods for humanoid sports have achieved success in single-agent and interaction-free settings, such as running, dancing, and parkour, task scenarios that require precise coordination among multiple participants remain largely unexplored. To this end, we propose Marope, a multi-agent reinforcement learning (MARL) framework for cooperative long rope skipping with multiple humanoid robots. Specifically, Marope adopts a hierarchical reinforcement learning framework for policy training. At the lower level, it learns decentralized rope manipulation policies through MARL, while at the upper level, a centralized scheduling policy is trained to coordinate the execution of the lower-level policies. To improve generalization across different player behavioral styles, Marope further incorporates diverse jumping policies into cooperative game training. We evaluate our approach on Unitree G1 humanoid robots in both simulation and real-world settings. Experimental results demonstrate that Marope outperforms various baselines, achieving more efficient and stable rope manipulation as well as more robust and adaptable cooperation with varied players.

多智能体强化学习人形机器人协作运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。