揭示了线性高斯老虎机中先验影响的独立作用机制。
Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
- 提出椭圆势能新工具,解耦先验与长期悔悟项。
- 证明贝叶斯悔悟上界为 $\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)})$。
- 首次证明先验烧入项不可避免,适用于贝叶斯强化学习研究者。
我们证明,在系数服从 $\mathcal{N}(μ_0, Σ_0)$ 先验的线性高斯老虎机中,Thompson采样具有 $\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)})$ 的贝叶斯悔悟。其中 $d$ 为维度,$T$ 为时间跨度,$r$ 为动作最大 $\ell_2$ 范数,$σ^2$ 为噪声方差。与以往结果不同,该界表明先验相关的‘烧入’项 $d r \sqrt{\mathrm{Tr}(Σ_0)}$ 在对数因子内可与最小最大(长期)悔悟项 $σd \sqrt{T}$ 加性分离,而非乘积依赖。通过新提出的‘椭圆势能’引理建立该结果,并给出下界说明烧入项不可避免。
原文摘要 · Abstract (English)
We prove that Thompson sampling exhibits $\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)})$ Bayesian regret in the linear-Gaussian bandit with a $\mathcal{N}(μ_0, Σ_0)$ prior distribution on the coefficients, where $d$ is the dimension, $T$ is the time horizon, $r$ is the maximum $\ell_2$ norm of the actions, and $σ^2$ is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ``burn-in'' term $d r \sqrt{\mathrm{Tr}(Σ_0)}$ decouples additively from the minimax (long run) regret $σd \sqrt{T}$. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ``elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。