arXiv:2512.24698cs.RO2025-12

用简化模型预训练+渐进迁移,让四足机器人更快学会翻滚等高动态动作

Dynamic Policy Learning for Legged Robot with Simplified Model Pretraining and Model-Homotopy-Inspired Transfer

  • 先用单刚体模型预训练,捕捉核心运动模式
  • 通过渐变质量分布路径,逐步迁移到真实四足模型,减少性能损失
  • 在真实机器人上成功实现翻滚、靠墙动作,收敛快且稳定

为腿式机器人生成动态运动仍具挑战。尽管强化学习在多种腿式运动任务中取得显著进展,但生成高度动态行为通常需大量奖励调参或高质量示范。利用简化模型可缓解此问题,但模型差异导致策略向全身体感环境迁移时面临挑战。本文提出一种基于连续性的学习框架,结合简化模型预训练与模型同伦启发的迁移策略,高效生成并优化复杂动态行为。首先,在单刚体模型环境中预训练策略以捕获核心运动模式;其次,采用连续性策略将策略逐步迁移至全身体感环境,最小化性能下降。为定义迁移路径,引入参数化过渡路径,逐步调整躯干与腿部的质量和惯性分布。该方法相比基线方法收敛更快,迁移过程更稳定。框架在翻滚、靠墙操作等多种动态任务中验证,并成功部署于真实四足机器人。

原文摘要 · Abstract (English)

Generating dynamic motions for legged robots remains a challenging problem. While reinforcement learning has achieved notable success in various legged locomotion tasks, producing highly dynamic behaviors often requires extensive reward tuning or high-quality demonstrations. Leveraging reduced-order models can help mitigate these challenges. However, the model discrepancy poses a significant challenge when transferring policies to full-body dynamics environments. In this work, we introduce a continuation-based learning framework that combines simplified model pretraining and model-homotopy-inspired transfer to efficiently generate and refine complex dynamic behaviors. First, we pretrain the policy using a single rigid body model to capture core motion patterns in a simplified environment. Next, we employ a continuation strategy to progressively transfer the policy to the full-body environment, minimizing performance loss. To define the continuation path, we introduce a parametric transition path from the single rigid body model to the full-body model by gradually redistributing mass and inertia between the trunk and legs. The proposed method achieves faster convergence and demonstrates superior stability during the transfer process compared to baseline methods. Our framework is validated on a range of dynamic tasks, including flips and wall-assisted maneuvers, and is successfully deployed on a real quadrupedal robot.

四足机器人强化学习动态运动迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。