提出新条件让贪心算法在多种分布下实现对数级累积后悔
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
- 引入局部反集中条件(LAC),使贪心算法可证明高效
- 在多种分布下,累积期望后悔上界为O(多项式对数T)
- 适用范围广,适合研究高效无探索算法的研究者
我们研究了线性上下文强化学习中无探索贪心算法的性能保证。提出一种新条件——局部反集中(Local Anti-Concentration, LAC)条件,使贪心算法能实现可证明的效率。我们证明该条件在广泛分布类中成立,包括高斯、指数、均匀、柯西、学生t分布以及其它指数族分布及其截断变体。在此条件下,贪心算法在线性上下文强化学习中的累积期望后悔被证明为O(poly log T)。本结果建立了迄今为止已知最广泛的允许贪心算法实现次线性后悔的分布范围,并达到尖锐的多项式对数后悔界。
原文摘要 · Abstract (English)
We study the performance guarantees of exploration-free greedy algorithms for the linear contextual bandit problem. We introduce a novel condition, named the \textit{Local Anti-Concentration} (LAC) condition, which enables a greedy bandit algorithm to achieve provable efficiency. We show that the LAC condition is satisfied by a broad class of distributions, including Gaussian, exponential, uniform, Cauchy, and Student's~$t$ distributions, along with other exponential family distributions and their truncated variants. This significantly expands the class of distributions under which greedy algorithms can perform efficiently. Under our proposed LAC condition, we prove that the cumulative expected regret of the greedy algorithm for the linear contextual bandit is bounded by $O(\operatorname{poly} \log T)$. Our results establish the widest range of distributions known to date that allow a sublinear regret bound for greedy algorithms, further achieving a sharp poly-logarithmic regret.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。