arXiv:2502.01226cs.LGstat.ML2025-02

自适应选择高斯过程先验,提升黑箱优化的准确性与效率

Adaptive Prior Selection in Gaussian Process Bandits with Thompson Sampling

  • 基于泰普森采样设计两种先验筛选算法,动态排除表现差的先验
  • 新算法在真实和合成数据上均显著降低累计损失(后悔值)
  • 适合需要自动调整先验的机器学习优化场景

高斯过程(GP)贝叶斯优化为未知函数的黑箱优化提供了强大框架,但其性能高度依赖于所假设的先验。现有研究通常假设先验已知,而实际中往往通过最大似然估计确定先验超参数,缺乏理论保障。本文提出两种基于高斯过程泰普森采样(GP-TS)的联合先验选择与后悔最小化算法:先验剔除型GP-TS(PE-GP-TS),通过预测性能淘汰劣质先验;超先验型GP-TS(HP-GP-TS),采用双层泰普森采样机制。理论上,我们建立了HP-GP-TS的次线性后悔界。实验表明,该方法在合成与真实数据上均优于基准方法。

原文摘要 · Abstract (English)

Gaussian process (GP) bandits provide a powerful framework for performing blackbox optimization of unknown functions. The characteristics of the unknown function depend heavily on the assumed GP prior. Most work in the literature assume that this prior is known but in practice this seldom holds. Instead, practitioners often rely on maximum likelihood estimation to select the hyperparameters of the prior - which lacks theoretical guarantees. In this work, we study two algorithms for joint prior selection and regret minimization in GP bandits based on GP Thompson sampling (GP-TS): Prior-Elimination GP-TS (PE-GP-TS) that disqualifies priors with poor predictive performance, and HyperPrior GP-TS (HP-GP-TS) that utilizes a bi-level Thompson sampling scheme. We theoretically analyze the algorithms and establish a sublinear regret bound for HP-GP-TS. In addition, we demonstrate the effectiveness of these algorithms compared to the alternatives through extensive experiments with synthetic and real-world data.

贝叶斯优化高斯过程强化学习自适应先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。