arXiv:2409.19617cs.RO2024-09中稿 · Robotics and Auton…被引 4

提出轻量鲁棒对抗学习框架,平衡机器人控制的稳健性与性能

LiRA: Light-Robust Adversary for Model-based Reinforcement Learning in Real World

  • 基于变分推断重构建模对抗学习,实现自适应对抗强度
  • 在性能下降可控范围内最大化鲁棒性,避免过度保守
  • 仅用2小时真实数据训练四足机器人抗力步态,效果显著

基于模型的强化学习因其高样本效率受到关注,有望应用于现实世界机器人。现实中不可观测干扰可能导致意外情况,因此机器人策略需兼顾控制性能与鲁棒性。对抗学习是提升鲁棒性的有效方法,但过度对抗会增加故障风险,使控制过于保守。为此,本文提出一种新的对抗学习框架——LiRA,通过变分推断重新构建对抗学习,并引入“轻量鲁棒性”约束:在可接受的性能退化范围内最大化鲁棒性。该框架可自动调节对抗强度,平衡鲁棒性与保守性。数值仿真验证了预期行为。此外,仅使用不到两小时的真实世界数据,LiRA成功训练出四足机器人的力反应步态控制策略。

原文摘要 · Abstract (English)

Model-based reinforcement learning has attracted much attention due to its high sample efficiency and is expected to be applied to real-world robotic applications. In the real world, as unobservable disturbances can lead to unexpected situations, robot policies should be taken to improve not only control performance but also robustness. Adversarial learning is an effective way to improve robustness, but excessive adversary would increase the risk of malfunction, and make the control performance too conservative. Therefore, this study addresses a new adversarial learning framework to make reinforcement learning robust moderately and not conservative too much. To this end, the adversarial learning is first rederived with variational inference. In addition, \textit{light robustness}, which allows for maximizing robustness within an acceptable performance degradation, is utilized as a constraint. As a result, the proposed framework, so-called LiRA, can automatically adjust adversary level, balancing robustness and conservativeness. The expected behaviors of LiRA are confirmed in numerical simulations. In addition, LiRA succeeds in learning a force-reactive gait control of a quadrupedal robot only with real-world data collected less than two hours.

强化学习机器人控制鲁棒性对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。