通过智能筛选攻击点,提升人形机器人长期运动的稳定性。
Keep on Going: Learning Robust Humanoid Motion Skills via Selective Adversarial Training
- 只对最脆弱状态和动作施加干扰,避免过度保守训练。
- 实测在复杂地形上成功率提升40%,轨迹误差降低32%。
- 适合需要长时间稳定运行的人形机器人研发团队。
人形机器人需在长时间运行中可靠执行全身动作,但强化学习运动策略常因传感器/执行器噪声及现实干扰而失稳。本文提出选择性对抗攻击鲁棒训练(SA2RT),让对抗方学习识别并稀疏扰动最脆弱的状态与动作,在攻击预算约束下暴露真实弱点,避免保守过拟合。非零和交替优化持续强化运动策略以应对最强攻击。我们在Unitree G1人形机器人上验证了该方法,涵盖感知行走与全身控制任务。实验表明,对抗训练后的策略使地形穿越成功率提升40%,轨迹跟踪误差降低32%,并保持长期移动与跟踪性能。结果表明,选择性对抗攻击是提升人形机器人长期运动鲁棒性的有效手段。
原文摘要 · Abstract (English)
Humanoid robots are expected to operate reliably over long horizons while executing versatile whole-body skills. Yet Reinforcement Learning (RL) motion policies typically lose stability under prolonged operation, sensor/actuator noise, and real world disturbances. In this work, we propose a Selective Adversarial Attack for Robust Training (SA2RT) to enhance the robustness of motion skills. The adversary is learned to identify and sparsely perturb the most vulnerable states and actions under an attack-budget constraint, thereby exposing true weakness without inducing conservative overfitting. The resulting non-zero sum, alternating optimization continually strengthens the motion policy against the strongest discovered attacks. We validate our approach on the Unitree G1 humanoid robot across perceptive locomotion and whole-body control tasks. Experimental results show that adversarially trained policies improve the terrain traversal success rate by 40%, reduce the trajectory tracking error by 32%, and maintain long horizon mobility and tracking performance. Together, these results demonstrate that selective adversarial attacks are an effective driver for learning robust, long horizon humanoid motion skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。