arXiv:2410.00859eess.SYcs.LG2024-10中稿 · publication in CDC…

用障碍函数改进强化学习中专家控制器的平滑性,提升样本效率。

Improved Sample Complexity of Imitation Learning for Barrier Model Predictive Control

  • 基于障碍函数重构MPC,生成平滑稳定的专家策略
  • 理论证明平滑性与误差权衡达到最优,优于随机平滑
  • 适用于带状态和输入约束的系统,适合控制类研究者

近期模仿学习研究表明,若专家控制器具备良好的平滑性和稳定性,则可对学习到的控制器性能提供更强保证。然而,在存在输入和状态约束的情况下,为任意系统构造此类平滑专家控制器仍具挑战。本文提出一种通用方法:通过将标准模型预测控制(MPC)优化问题转化为基于对数障碍函数的松弛形式,设计出满足要求的平滑专家控制器。相比先前工作,本文证明障碍型MPC在特定方向上实现了理论上最优的误差-平滑性权衡。该理论保障的核心在于我们对凸Lipschitz函数对应解析中心的最优性间隙给出了更优下界,这一结果或具有独立研究价值。实验验证了理论发现,表明所提平滑方法在性能上显著优于随机平滑。

原文摘要 · Abstract (English)

Recent work in imitation learning has shown that having an expert controller that is both suitably smooth and stable enables stronger guarantees on the performance of the learned controller. However, constructing such smoothed expert controllers for arbitrary systems remains challenging, especially in the presence of input and state constraints. As our primary contribution, we show how such a smoothed expert can be designed for a general class of systems using a log-barrier-based relaxation of a standard Model Predictive Control (MPC) optimization problem. Improving upon our previous work, we show that barrier MPC achieves theoretically optimal error-to-smoothness tradeoff along some direction. At the core of this theoretical guarantee on smoothness is an improved lower bound we prove on the optimality gap of the analytic center associated with a convex Lipschitz function, which we believe could be of independent interest. We validate our theoretical findings via experiments, demonstrating the merits of our smoothing approach over randomized smoothing.

模仿学习MPC平滑性障碍函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。