解决模仿学习中约束不可行问题,提升训练稳定性。
Infeasible optimization problems and the hierarchical augmented Lagrangian method in imitation learning

- 引入分层增广拉格朗日法处理不可行约束
- 使策略逼近最接近可行的约束优化解
- 适用于需安全约束的机器人控制场景
模仿学习(IL)是训练复杂机器人策略的有效方法。近期研究将硬约束引入模仿学习优化问题,以保障所学策略的安全性、稳定性和鲁棒性。然而,我们指出这些约束有时可能不可行,导致训练动态不稳定或难以收敛。本文基于增广拉格朗日法在不可行情形下的最新理论,提出一种简单修复方案。该方法可引导学习策略趋近于一个具有理想性质的最近可行约束模仿学习问题的解。我们在一个包含总加速度约束和行人安全约束的模拟驾驶任务中验证了该方法,该设置下不可行性可能自然出现,但仍能学习到安全策略。
原文摘要 · Abstract (English)
Imitation learning (IL) is an effective approach to train complex robotics policies. Recent works have introduced hard constraints into imitation-learning optimization problems to ensure safety, stability, and robustness of the learned policy. However, we argue that these constraints are sometimes infeasible, which can lead to unstable or difficult training dynamics. We study a simple remedy for such situations based on recent theoretical results on the augmented Lagrangian method in infeasible settings. We show that our approach drives the learned policy toward the solution of a closest-feasible constrained IL problem with desirable properties. The method is illustrated on a toy driving example with a total-acceleration constraint and pedestrian-safety constraints, a setting in which infeasibility can naturally arise while still allowing a safe learned policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。