提出可自适应估计核空间范数的安全贝叶斯优化方法
Safe exploration in reproducing kernel Hilbert spaces
- 从数据中学习未知函数的再生核希尔伯特空间范数
- 在物理仿真与真实倒立摆上实现更安全高效的策略优化
- 适合需要动态安全约束的强化学习系统部署
主流安全贝叶斯优化(BO)算法用于在未知环境中学习安全控制策略,但多数依赖已知的再生核希尔伯特空间(RKHS)范数有界假设。由于RKHS是潜在无穷维空间,如何可靠估计未知函数的范数仍不明确。本文提出一种可自适应估计RKHS范数的安全BO算法,提供该估计的统计保证,并将其融入现有置信区间,保持理论安全性。在物理模拟器和真实倒立摆系统上验证,相比当前最优方法,本算法在性能、安全性和可扩展性上均有提升。
原文摘要 · Abstract (English)
Popular safe Bayesian optimization (BO) algorithms learn control policies for safety-critical systems in unknown environments. However, most algorithms make a smoothness assumption, which is encoded by a known bounded norm in a reproducing kernel Hilbert space (RKHS). The RKHS is a potentially infinite-dimensional space, and it remains unclear how to reliably obtain the RKHS norm of an unknown function. In this work, we propose a safe BO algorithm capable of estimating the RKHS norm from data. We provide statistical guarantees on the RKHS norm estimation, integrate the estimated RKHS norm into existing confidence intervals and show that we retain theoretical guarantees, and prove safety of the resulting safe BO algorithm. We apply our algorithm to safely optimize reinforcement learning policies on physics simulators and on a real inverted pendulum, demonstrating improved performance, safety, and scalability compared to the state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。