arXiv:2608.18863stat.MLcs.LG2026-08

常数探索下实现更优后悔率,适用于动态环境的贝叶斯优化。

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

  • 用局部置信事件替代全局约束,允许恒定探索参数。
  • 在持续漂移下,平均后悔率达到$ ilde{/mathcal O}(ε^{1/4})$。
  • 适合需要稳定探索策略的实时优化场景。

我们研究了在随时间变化环境中基于高斯过程漂移模型的贝叶斯优化问题。现有GP-UCB分析通常需随时间增长探索参数以维持一致置信边界。本文通过每轮局部置信事件,证明可使用恒定探索参数,并获得依赖漂移速率的期望后悔率上界。同时推导出更精细的时间变最大信息增益界。对于平方指数核,在持续漂移情形下,$ ildeγ_T/T= ilde{/mathcal O}(ε^{1/2})$,期望平均后悔率$ ilde{/mathcal O}(ε^{1/4})$。相同分析也给出实现后悔率保证。模拟验证了边界建议的探索参数对$1/ε$的对数依赖关系。

原文摘要 · Abstract (English)

We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence events, we show that GP-UCB can instead be run with a constant exploration parameter and obtain an expected-regret bound whose coefficient depends on the drift rate. We also derive a sharper time-varying maximum-information-gain bound. For the squared exponential kernel, it yields $\tildeγ_T/T=\widetilde{\mathcal O}(ε^{1/2})$ and expected average regret $\widetilde{\mathcal O}(ε^{1/4})$ in the persistent-drift regime. The same constant-exploration analysis also yields realized-regret guarantees. Simulations support the predicted logarithmic dependence of the bound-suggested exploration parameter on $1/ε$.

贝叶斯优化高斯过程动态环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。