解决动态函数优化的无遗憾问题,通过引入不确定性提升模型适应性。
No-Regret Gaussian Process Optimization of Time-Varying Functions
- 用不确定性注入实现时变函数的异方差高斯过程回归
- 在允许少量历史点重查条件下达到无遗憾性能
- 理论证明方法高效,适用于随时间变化的优化场景
从噪声观测中顺序优化黑箱函数广泛研究,传统高斯过程强化学习算法如GP-UCB在静态设定下可保证无遗憾。然而对于时变目标,在纯强化学习反馈下,除非施加强且不现实的假设,否则无法实现无遗憾。本文提出一种新方法,在频率论设定下优化具有有界再生核希尔伯特空间(RKHS)范数的时间变化奖励函数。通过不确定性注入捕捉时变特性,实现自适应过去观测的异方差高斯过程回归。由于严格强化学习设定下无遗憾不可达,我们放宽限制,允许对先前观测点进行额外查询。基于稀疏推理与不确定性注入对遗憾的影响,提出W-SparQ-GP-UCB算法,可在每轮迭代中以趋于零的额外查询数实现无遗憾。为评估该方法的理论极限,建立了实现无遗憾所需的额外查询数下界,证明了方法效率。最后,全面分析了函数时间演化模式与可实现遗憾率的关系,并给出各演化阶段所需额外查询数的上下界。
原文摘要 · Abstract (English)
Sequential optimization of black-box functions from noisy evaluations has been widely studied, with Gaussian Process bandit algorithms such as GP-UCB guaranteeing no-regret in stationary settings. However, for time-varying objectives, no-regret is unattainable under pure bandit feedback unless strong and often unrealistic assumptions are imposed. We propose a novel method for optimizing time-varying rewards in the frequentist setting, where the objective has bounded RKHS norm almost surely. Time variations are captured through uncertainty injection, enabling heteroscedastic Gaussian process regression that adapts past observations to the current time step. As no-regret is unattainable in general in the strict bandit setting, we relax the latter allowing additional queries on previously observed points. Building on sparse inference and the effect of uncertainty injection on regret, we propose W-SparQ-GP-UCB, an online algorithm that achieves no-regret with a vanishing number of additional queries per iteration. To assess the theoretical limits of this approach, we establish a lower bound on the number of additional queries required for no-regret, proving the efficiency of our method. Finally, we provide a comprehensive analysis linking the temporal regime of the function to achievable regret rates, together with upper and lower bounds on the number of additional queries needed in each regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。