arXiv:2512.12230cs.ROcs.LG2025-12中稿 · 28th RoboCup Inter…被引 3

一个策略搞定七种人形机器人跌倒恢复,无需重新训练

Learning to Get Up Across Morphologies: Zero-Shot Recovery with a Unified Humanoid Policy

  • 用统一的强化学习策略跨形态实现跌倒起身
  • 零样本迁移成功率86%±7%,覆盖0.48~0.81米身高差异
  • 适合需快速部署、多型号通用的人形机器人应用

在动态环境如RoboCup中,跌倒恢复能力对人形机器人的表现至关重要,长时间停机常导致比赛失利。现有基于深度强化学习(DRL)的方法虽能生成稳健的起身行为,但需为每种机器人构型单独训练策略。本文提出一种单一DRL策略,可在七种不同构型的人形机器人上实现跌倒恢复,其高度范围为0.48–0.81米,重量范围为2.8–7.9千克,动力学特性各异。该策略通过CrossQ训练,在未见过的构型上实现86%±7%(95%置信区间[81, 89])的零样本迁移成功率,无需针对特定机器人训练。通过留一法实验、构型缩放分析及多样性消融实验表明,针对性地覆盖构型差异可显著提升零样本泛化能力。在某些情况下,共享策略甚至优于专用基线。这些结果验证了形态无关控制在跌倒恢复中的实用性,为人形机器人通用控制奠定了基础。代码已开源:https://github.com/utra-robosoccer/unified-humanoid-getup

原文摘要 · Abstract (English)

Fall recovery is a critical skill for humanoid robots in dynamic environments such as RoboCup, where prolonged downtime often decides the match. Recent techniques using deep reinforcement learning (DRL) have produced robust get-up behaviors, yet existing methods require training of separate policies for each robot morphology. This paper presents a single DRL policy capable of recovering from falls across seven humanoid robots with diverse heights (0.48 - 0.81 m), weights (2.8 - 7.9 kg), and dynamics. Trained with CrossQ, the unified policy transfers zero-shot up to 86 +/- 7% (95% CI [81, 89]) on unseen morphologies, eliminating the need for robot-specific training. Comprehensive leave-one-out experiments, morph scaling analysis, and diversity ablations show that targeted morphological coverage improves zero-shot generalization. In some cases, the shared policy even surpasses the specialist baselines. These findings illustrate the practicality of morphology-agnostic control for fall recovery, laying the foundation for generalist humanoid control. The software is open-source and available at: https://github.com/utra-robosoccer/unified-humanoid-getup

人形机器人跌倒恢复零样本迁移强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。