arXiv:2503.06226eess.SYcs.AI2025-03被引 1

无需状态观测收敛即可实现最优输出反馈控制,解决未知离散系统LQR难题。

Optimal Output Feedback Learning Control for Discrete-Time Linear Quadratic Regulation

  • 通过动态输出反馈等效于状态反馈,避免依赖观测器收敛
  • 基于价值迭代与策略迭代,自适应求解最优控制增益
  • 提出无模型稳定性判据,支持切换迭代机制,适合控制理论研究者

本文研究未知离散时间系统的线性二次调节(LQR)问题,采用动态输出反馈学习控制方法。与状态反馈不同,动态输出反馈的最优性依赖于状态观测器的隐式收敛条件,而系统矩阵未知和观测误差的存在使现有方法难以分析收敛性与稳定性。为此,提出一种广义动态输出反馈学习控制方法,确保收敛性、稳定性和最优性。关键在于设计的动态输出反馈控制器等价于状态反馈控制器,该等价关系不依赖观测器对状态的估计收敛,为非策略学习控制提供基础。基于价值迭代与策略迭代,发展了基于自适应动态规划的学习控制方法以估计最优反馈控制增益。此外,通过寻找非奇异参数化矩阵,提出无模型稳定性判据,支持切换迭代方案。最后,给出了所提方法的收敛性、稳定性和最优性分析,并通过两个数值例子验证理论结果。

原文摘要 · Abstract (English)

This paper studies the linear quadratic regulation (LQR) problem of unknown discrete-time systems via dynamic output feedback learning control. In contrast to the state feedback, the optimality of the dynamic output feedback control for solving the LQR problem requires an implicit condition on the convergence of the state observer. Moreover, due to unknown system matrices and the existence of observer error, it is difficult to analyze the convergence and stability of most existing output feedback learning-based control methods. To tackle these issues, we propose a generalized dynamic output feedback learning control approach with guaranteed convergence, stability, and optimality performance for solving the LQR problem of unknown discrete-time linear systems. In particular, a dynamic output feedback controller is designed to be equivalent to a state feedback controller. This equivalence relationship is an inherent property without requiring convergence of the estimated state by the state observer, which plays a key role in establishing the off-policy learning control approaches. By value iteration and policy iteration schemes, the adaptive dynamic programming based learning control approaches are developed to estimate the optimal feedback control gain. In addition, a model-free stability criterion is provided by finding a nonsingular parameterization matrix, which contributes to establishing a switched iteration scheme. Furthermore, the convergence, stability, and optimality analyses of the proposed output feedback learning control approaches are given. Finally, the theoretical results are validated by two numerical examples.

LQR输出反馈自适应控制学习控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。