arXiv:2412.02423eess.SYcs.LG2024-12被引 3

利用时间序列结构加速闭环控制参数优化,节省实验资源。

Time-Series-Informed Closed-loop Learning for Sequential Decision Making and Control

  • 将多保真度贝叶斯优化与闭环时间对齐,用中间结果提升效率。
  • 实验资源减半,性能相当;相同资源下表现更优。
  • 适合需要高效调参的机器人、工业控制等闭环系统场景。

闭环决策算法(如模型预测控制)的性能高度依赖控制器参数选择。标准贝叶斯优化将此视为黑箱问题,忽略闭环轨迹的时间结构,导致收敛慢、资源浪费。本文提出一种时间序列感知的多保真度贝叶斯优化框架,将保真度维度与闭环时间对齐,使实验中的中间性能评估可作为低保真观测纳入。同时,推导了基于代理模型后验信念的概率早期停止准则,提前终止无望的实验,避免完整运行差参数配置,降低资源消耗。在非线性控制基准测试中,相比传统黑箱方法,本方法以约一半实验资源达成相当性能,相同资源下获得更好最终性能,证明利用时间结构对样本高效闭环调参具有显著价值。

原文摘要 · Abstract (English)

Closed-loop performance of sequential decision making algorithms, such as model predictive control, depends strongly on the choice of controller parameters. Bayesian optimization allows learning of parameters from closed-loop experiments, but standard Bayesian optimization treats this as a black-box problem and ignores the temporal structure of closed-loop trajectories, leading to slow convergence and inefficient use of experimental resources. We propose a time-series-informed multi-fidelity Bayesian optimization framework that aligns the fidelity dimension with closed-loop time, enabling intermediate performance evaluations within a closed-loop experiment to be incorporated as lower-fidelity observations. Additionally, we derive probabilistic early stopping criteria to terminate unpromising closed-loop experiments based on the surrogate model's posterior belief, avoiding full episodes for poor parameterizations and thereby reducing resource usage. Simulation results on a nonlinear control benchmark demonstrate that, compared to standard black-box Bayesian optimization approaches, the proposed method achieves comparable closed-loop performance with roughly half the experimental resources, and yields better final performance when using the same resource budget, highlighting the value of exploiting temporal structure for sample-efficient closed-loop controller tuning.

闭环控制贝叶斯优化多保真度参数调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。