用动态特征对抗训练,让人形机器人更稳地走远路。
ADP: Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion

- 用运动轨迹优化构建参考数据,以动态特征为对抗目标
- 抗扰能力提升16.7%,恢复时间减少47.9%
- 适合需要强鲁棒性的机器人行走控制场景
本文提出对抗性动态先验(ADP),用于实现抗扰的人形机器人行走控制。现有基于运动先验的方法通过模仿运动学特征生成自然动作,但未直接正则化质心运动、重心动量、接触力和接触状态等动力学特征。为此,我们以从运动轨迹中提取的动力学特征作为对抗正则化的目标。通过轨迹优化构建参考数据集,并训练判别器评估策略生成的时间窗口是否与参考分布一致。无需显式运动追踪,ADP 能使策略在扰动后仍保持在参考支撑范围内。实验表明,相比最强基线AMP,ADP将80%成功率冲击阈值(J₈₀)提高16.7%,方向平均恢复时间减少47.9%,速度跟踪误差降低35.4%。
原文摘要 · Abstract (English)
In this paper, we propose Adversarial Dynamics Priors (ADP) for perturbation-resilient humanoid locomotion control. Existing motion prior-based methods induce natural motion styles by imitating kinematic motion features, but they do not directly regularize dynamics features, such as CoM motion, centroidal momentum, contact forces, and contact states. To address this limitation, we replace kinematic motion-style feature with selected dynamics features extracted from locomotion trajectories as the target of adversarial regularization. To this end, we use trajectory optimization to construct a reference dataset and train a discriminator to evaluate whether policy-induced temporal windows are consistent with the resulting reference distribution. Without explicit motion tracking, ADP encourages policy rollouts to remain close to the reference support, even after perturbations. Experimental results show that, compared with AMP, the strongest baseline in our evaluation, ADP improves the $80\%$-success impulse threshold ($J_{80}$) by $16.7\%$, while reducing direction-averaged recovery time and velocity tracking error by $47.9\%$ and $35.4\%$, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。