arXiv:2506.22186cs.LG2025-06

用贝叶斯采样在线学习未知非线性系统的控制律,兼顾探索与效率。

Thompson Sampling-Based Learning and Control for Unknown Dynamic Systems

  • 将控制律视为函数空间中的元素,不依赖具体系统结构
  • 理论证明学习速度呈指数级,控制损失有上界
  • 适合对未知动态系统进行数据驱动的智能控制设计

Thompson采样(TS)是一种基于贝叶斯的随机探索策略,通过从当前后验中采样系统参数或控制律,并选择对任务最优的选项,实现探索与利用的平衡,适用于基于主动学习的控制器设计。然而,传统TS依赖有限参数表示,限制了其在更广泛控制空间中的应用。为此,本文提出一种基于再生核希尔伯特空间(RKHS)的控制律参数化方法,构建数据驱动的主动学习控制框架。该方法将控制律视为函数空间中的元素,无需预设系统结构或控制器形式。基于TS框架,实现在线探索与利用以降低控制成本,并提供学习过程的收敛性保证。理论分析表明,所提方法可使控制律与闭环性能之间的关系以指数速率学习,且推导出控制遗憾的上界。此外,还分析了闭环稳定性。在未知非线性系统的数值实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Thompson sampling (TS) is a Bayesian randomized exploration strategy that samples options (e.g., system parameters or control laws) from the current posterior and then applies the selected option that is optimal for a task, thereby balancing exploration and exploitation; this makes TS effective for active learning-based controller design. However, TS relies on finite parametric representations, which limits its applicability to more general spaces, which are more commonly encountered in control system design. To address this issue, this work proposes a parameterization method for control law learning using reproducing kernel Hilbert spaces and designs a data-driven active learning control approach. Specifically, the proposed method treats the control law as an element in a function space, allowing the design of control laws without imposing restrictions on the system structure or the form of the controller. A TS framework is proposed in this work to reduce control costs through online exploration and exploitation, and the convergence guarantees are further provided for the learning process. Theoretical analysis shows that the proposed method learns the relationship between control laws and closed-loop performance metrics at an exponential rate, and the upper bound of control regret is also derived. Furthermore, the closed-loop stability of the proposed learning framework is analyzed. Numerical experiments on controlling unknown nonlinear systems validate the effectiveness of the proposed method.

控制学习贝叶斯优化非线性控制主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。