arXiv:2603.28938eess.SYcs.LG2026-03被引 1

用内在奖励提升探索效率,实现最优在线控制。

Optimistic Online LQR via Intrinsic Rewards

  • 引入内在奖励与方差正则化,驱动不确定性下的探索
  • 达到理论最优的√T最坏情况后悔率
  • 结构简单高效,适合实时控制场景

面对未知线性动态系统时,在线线性二次调节器(LQR)问题要求基于运行中收集的闭环数据,自适应调整控制策略。本文提出内在奖励LQR(IR-LQR),一种乐观在线LQR算法,结合强化学习中的内在奖励思想与方差正则化机制,促进由不确定性驱动的探索。IR-LQR仅通过修改代价函数,保持标准LQR求解结构,具有直观、简洁、计算成本低、高效等优势。相比现有依赖复杂迭代搜索或求解高计算开销优化问题的乐观在线LQR方法,本方法更具实用性。理论上,IR-LQR实现了最优的最坏情况后悔率√T。在飞机俯仰角控制与无人飞行器实例上的数值实验表明,其性能优于多种前沿在线LQR算法。

原文摘要 · Abstract (English)

Optimism in the face of uncertainty is a popular approach to balance exploration and exploitation in reinforcement learning. Here, we consider the online linear quadratic regulator (LQR) problem, i.e., to learn the LQR corresponding to an unknown linear dynamical system by adapting the control policy online based on closed-loop data collected during operation. In this work, we propose Intrinsic Rewards LQR (IR-LQR), an optimistic online LQR algorithm that applies the idea of intrinsic rewards originating from reinforcement learning and the concept of variance regularization to promote uncertainty-driven exploration. IR-LQR typically retains the structure of a standard LQR synthesis problem by only modifying the cost function, resulting in an intuitively pleasing, simple, computationally cheap, and efficient algorithm. This is in contrast to existing optimistic online LQR formulations that rely on more complicated iterative search algorithms or solve computationally demanding optimization problems. We show that IR-LQR achieves the optimal worst-case regret rate of $\sqrt{T}$, and compare it to various state-of-the-art online LQR algorithms via numerical experiments carried out on an aircraft pitch angle control and an unmanned aerial vehicle example.

在线控制强化学习最优控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。