arXiv:2602.14472math.STcs.LG2026-02

提出无需离散化的新分析框架,统一证明高斯过程强化学习的后悔界。

Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors

  • 用分数后验解释常见方差膨胀,构建新分析视角。
  • 给出与核函数无关的后悔上界,涵盖平方指数、Matérn等核。
  • 适用于连续动作空间,为算法设计提供理论支撑。

我们研究在紧致连续动作空间上的高斯过程汤普森采样(GP-TS),并基于分数高斯过程后验提供非贝叶斯后悔分析,无需依赖先前工作中的域离散化。我们发现现有分析中常见的方差膨胀可被解释为以分数后验(温度参数α∈(0,1))进行汤普森采样。我们推导出一个不依赖核函数的后悔上界,表达为信息增益参数γ_t和后验收缩率ε_t的函数,并识别出在何种先验条件下可控制ε_t。作为特例,我们恢复了平方指数核下的 ilde{/mathcal{O}}(T^{1/2})后悔界,以及Matérn-ν核和有理二次核下的 ilde{/mathcal{O}}(T^{(2ν+3d)/(2(2ν+d))})后悔界。总体而言,该分析为GP-TS提供了统一且无需离散化的后悔框架,适用于多种核类型。

原文摘要 · Abstract (English)

We study Gaussian Process Thompson Sampling (GP-TS) for sequential decision-making over compact, continuous action spaces and provide a frequentist regret analysis based on fractional Gaussian process posteriors, without relying on domain discretization as in prior work. We show that the variance inflation commonly assumed in existing analyses of GP-TS can be interpreted as Thompson Sampling with respect to a fractional posterior with tempering parameter $α\in (0,1)$. We derive a kernel-agnostic regret bound expressed in terms of the information gain parameter $γ_t$ and the posterior contraction rate $ε_t$, and identify conditions on the Gaussian process prior under which $ε_t$ can be controlled. As special cases of our general bound, we recover regret of order $\tilde{\mathcal{O}}(T^{\frac{1}{2}})$ for the squared exponential kernel, $\tilde{\mathcal{O}}(T^{\frac{2ν+3d}{2(2ν+d)}} )$ for the Matérn-$ν$ kernel, and a bound of order $\tilde{\mathcal{O}}(T^{\frac{2ν+3d}{2(2ν+d)}})$ for the rational quadratic kernel. Overall, our analysis provides a unified and discretization-free regret framework for GP-TS that applies broadly across kernel classes.

强化学习高斯过程后悔分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。