arXiv:2509.24539cs.RO2025-09

用SAC替代PPO,让机器人模仿动物运动更高效稳定。

Unlocking the Potential of Soft Actor-Critic for Imitation Learning

  • 将对抗性运动先验与离策略SAC结合,提升数据利用效率。
  • 在多种地形上实现更高模仿奖励和稳定任务执行。
  • 适合追求高效、泛化强的机器人运动生成研究者。

基于学习的方法使机器人能够逐步掌握类生物运动,具备更高的自然性和适应性。其中,模仿学习(IL)在将动物复杂运动模式迁移到机器人系统方面已证明有效。然而,当前主流框架多依赖于对策略算法PPO,其虽具稳定性,但牺牲了样本效率与策略泛化能力。本文提出一种新型模仿学习框架,将对抗性运动先验(AMP)与离策略的软演员-评论家(SAC)算法相结合,利用回放缓冲区学习和熵正则化探索,实现自然行为与任务执行,显著提升数据效率与鲁棒性。我们在涉及多种参考运动和多样地形的四足步态上评估该方法(AMP+SAC)。实验结果表明,该框架不仅保持任务执行稳定,且在模仿奖励上优于广泛使用的AMP+PPO方法。这些发现凸显了离策略模仿学习在机器人运动生成中的潜力。

原文摘要 · Abstract (English)

Learning-based methods have enabled robots to acquire bio-inspired movements with increasing levels of naturalness and adaptability. Among these, Imitation Learning (IL) has proven effective in transferring complex motion patterns from animals to robotic systems. However, current state-of-the-art frameworks predominantly rely on Proximal Policy Optimization (PPO), an on-policy algorithm that prioritizes stability over sample efficiency and policy generalization. This paper proposes a novel IL framework that combines Adversarial Motion Priors (AMP) with the off-policy Soft Actor-Critic (SAC) algorithm to overcome these limitations. This integration leverages replay-driven learning and entropy-regularized exploration, enabling naturalistic behavior and task execution, improving data efficiency and robustness. We evaluate the proposed approach (AMP+SAC) on quadruped gaits involving multiple reference motions and diverse terrains. Experimental results demonstrate that the proposed framework not only maintains stable task execution but also achieves higher imitation rewards compared to the widely used AMP+PPO method. These findings highlight the potential of an off-policy IL formulation for advancing motion generation in robotics.

模仿学习强化学习机器人运动SAC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。