用贝叶斯优化安全训练神经网络MPC,保证系统长期稳定
Safe and Stable Closed-Loop Learning for Neural-Network-Supported Model Predictive Control
- 用神经网络参数化MPC阶段代价函数,提升控制灵活性
- 在闭环数据中优化参数,实现长期性能与稳定性的平衡
- 将稳定性信息嵌入贝叶斯优化,获得概率性安全保证
安全学习控制策略在最优控制和强化学习中仍具挑战性。本文研究在对底层过程信息不完全的情况下,参数化预测控制器的安全学习问题。为此,我们采用贝叶斯优化从闭环数据中学习最优参数。该方法关注系统在闭环下的整体长期性能,同时确保其安全与稳定。具体而言,我们使用前馈神经网络参数化模型预测控制(MPC)的阶段代价函数,从而实现高灵活性,使系统在更高层次指标下获得更优的闭环性能。然而,这种灵活性也要求严格的安全部署,特别是闭环稳定性保障。为此,我们在基于贝叶斯优化的学习过程中显式引入稳定性信息,实现了严格的概率性安全保证。该方法通过数值案例进行了验证。
原文摘要 · Abstract (English)
Safe learning of control policies remains challenging, both in optimal control and reinforcement learning. In this article, we consider safe learning of parametrized predictive controllers that operate with incomplete information about the underlying process. To this end, we employ Bayesian optimization for learning the best parameters from closed-loop data. Our method focuses on the system's overall long-term performance in closed-loop while keeping it safe and stable. Specifically, we parametrize the stage cost function of an MPC using a feedforward neural network. This allows for a high degree of flexibility, enabling the system to achieve a better closed-loop performance with respect to a superordinate measure. However, this flexibility also necessitates safety measures, especially with respect to closed-loop stability. To this end, we explicitly incorporated stability information in the Bayesian-optimization-based learning procedure, thereby achieving rigorous probabilistic safety guarantees. The proposed approach is illustrated using a numeric example.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。