统一训练人形机器人运动追踪与跌倒恢复,提升鲁棒性。
Stubborn: A Streamlined and Unified Reinforcement Learning Framework for Robust Motion Tracking and Fall Recovery for Humanoids

- 采用非对称演员-评论家架构,结合航向对齐表征减少漂移影响。
- 引入贝努利概率终止机制,支持在跌倒状态持续探索恢复动作。
- 动态调整采样分布,增强复杂动作段和不稳定状态的训练效率。
近期强化学习方法在提升人形机器人运动追踪性能和实现扰动下的跌倒恢复方面展现出巨大潜力。然而,多数现有工作将运动追踪与跌倒恢复视为独立任务,需分阶段训练并依赖专用恢复奖励或独立恢复策略。此外,现有基于强化学习的方法通常在严重追踪失败后立即终止训练,限制了在不稳或倒地状态下的恢复探索。为解决上述问题,我们提出 Stubborn——一种简化的统一强化学习框架,用于实现鲁棒的人形机器人运动追踪与跌倒恢复。具体而言,Stubborn 采用非对称演员-评论家架构,包含三个核心组件:首先,引入航向对齐的追踪表示,降低对全局漂移和航向扰动的敏感性,同时保留重力相关的平衡信息;其次,提出基于贝努利的概率终止机制,使策略能够在不同故障模式下鼓励跌倒恢复行为的探索;第三,设计基于概率终止与追踪误差驱动的策略,动态调整采样分布,提升困难运动片段及不稳定状态的训练效率。大量与最先进方法的对比实验及消融研究显示,Stubborn 实现了具有竞争力的性能,所提出的概率终止机制与自适应采样策略显著提升了性能与鲁棒性。真实世界演示请见 https://aislab-sustech.github.io/Stubborn/
原文摘要 · Abstract (English)
Recent reinforcement learning approaches have shown great promise in improving humanoid motion tracking performance and achieving fall recovery under disturbances. However, most existing works treat motion tracking and fall recovery as different tasks and require multi-stage training with specialized recovery rewards and/or separate recovery policies. Moreover, existing reinforcement learning-based methods often terminate training episodes immediately after severe tracking failures, limiting recovery-oriented exploration in unstable or fallen states. To address the above issues, we propose Stubborn, a streamlined and unified reinforcement learning framework to achieve robust humanoid motion tracking and fall recovery. Specifically, Stubborn uses an asymmetric Actor-Critic architecture and consists of three major components. First, a yaw-aligned tracking representation is adopted to reduce sensitivity to global drift and heading disturbances while preserving gravity-related balance information. Second, we introduce a Bernoulli-based probabilistic termination mechanism that enables the policy to encourage exploration of fall-recovery behaviors under varying failure modes. Third, we propose a probabilistic termination and tracking-error-driven strategy that dynamically reshapes the sampling distribution based on tracking performance, increasing the training efficiency for difficult motion segments and unstable states. Extensive comparisons with SOTA methods and ablation studies show that Stubborn achieved competitive performance, and the proposed probabilistic termination mechanism and adaptive sampling strategy contributed to the performance and robustness gains. For real-world demonstrations, please refer to https://aislab-sustech.github.io/Stubborn/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。