arXiv:2505.13549cs.RO2025-05被引 2

提出新方法让人形机器人走路更稳,无需改底层规划器。

TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion

  • 用隐空间信任域约束保持规划与真实动作一致
  • 在26自由度人形机上实现复杂动态行走,训练更高效
  • 适合做高维控制的机器人学习研究者

在高维控制场景(如人形机器人行走)中,强化学习面临动态不稳定、接触交互复杂及训练分布偏移敏感等问题。基于模型的方法如时序差分模型预测控制(TD-MPC)结合短期规划与价值学习,在基础行走任务中表现良好。但现有方法难以解决离策略更新带来的策略不匹配与不稳定问题。为此,本文提出时序差分组相对策略约束(TD-GRPC),在TD-MPC框架下融合组相对策略优化(GRPO)与显式策略约束(PC)。TD-GRPC在潜在策略空间施加信任域约束,维持规划先验与学习轨迹的一致性,并利用组相对排序评估并保留候选轨迹的物理可行性。该方法无需修改底层规划器,即可实现灵活规划与策略学习。我们在26自由度的Unitree H1-2人形机器人上验证了该方法在从基础行走到高度动态运动的任务套件中的有效性。仿真结果表明,TD-GRPC在复杂人形控制任务中提升了稳定性与策略鲁棒性,同时保持采样效率。

原文摘要 · Abstract (English)

Robot learning in high-dimensional control settings, such as humanoid locomotion, presents persistent challenges for reinforcement learning (RL) algorithms due to unstable dynamics, complex contact interactions, and sensitivity to distributional shifts during training. Model-based methods, \textit{e.g.}, Temporal-Difference Model Predictive Control (TD-MPC), have demonstrated promising results by combining short-horizon planning with value-based learning, enabling efficient solutions for basic locomotion tasks. However, these approaches remain ineffective in addressing policy mismatch and instability introduced by off-policy updates. Thus, in this work, we introduce Temporal-Difference Group Relative Policy Constraint (TD-GRPC), an extension of the TD-MPC framework that unifies Group Relative Policy Optimization (GRPO) with explicit Policy Constraints (PC). TD-GRPC applies a trust-region constraint in the latent policy space to maintain consistency between the planning priors and learned rollouts, while leveraging group-relative ranking to assess and preserve the physical feasibility of candidate trajectories. Unlike prior methods, TD-GRPC achieves robust motions without modifying the underlying planner, enabling flexible planning and policy learning. We validate our method across a locomotion task suite ranging from basic walking to highly dynamic movements on the 26-DoF Unitree H1-2 humanoid robot. Through simulation results, TD-GRPC demonstrates its improvements in stability and policy robustness with sampling efficiency while training for complex humanoid control tasks.

机器人控制强化学习人形机器人策略约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。