提出新算法,让非平稳线性博弈中找最优选项更高效。
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
- 基于邻近最优设计,动态调整采样策略。
- 误差概率逼近理论下界,性能随臂集结构优化。
- 适合有几何结构的复杂决策场景使用。
研究非平稳线性多臂老虎机中的固定预算最优臂识别问题。给定时间预算 $T$、有限臂集 $/mathcal{X} ackslashsubset /mathbb{R}^d$ 及潜在对抗性未知参数序列 $\lbrace θ_t brace_{t=1}^{T}$,学习者需以高概率识别出累积奖励最大之臂 $x_* = \ ext{argmax}_{x \\in \\mathcal{X}} x^ op\\sum_{t=1}^T θ_t$。现有方法在标准基向量臂集下达到 $\ ext{exp}(-Θ(T / H_G))$ 的最坏情况误差率,其中 $H_G \propto d$。但此下界过于悲观,未考虑臂集几何结构优势。本文建立依赖臂集的下界,并提出邻近最优设计与 $\ extsf{Adjacent-BAI}$ 算法,其误差率匹配下界,证明了复杂度与臂集结构相关,实现理论最优。
原文摘要 · Abstract (English)
We study the fixed-budget best-arm identification (BAI) problem in non-stationary linear bandits. Concretely, given a fixed time budget $T\in \mathbb{N}$, finite arm set $\mathcal{X} \subset \mathbb{R}^d$, and a potentially adversarial sequence of unknown parameters $\lbrace θ_t\rbrace_{t=1}^{T}$ (hence non-stationary), a learner aims to identify the arm with the largest cumulative reward $x_* = \arg\max_{x \in \mathcal{X}} x^\top\sum_{t=1}^T θ_t$ with high probability. In this setting, it is well-known that uniformly sampling arms from the G-optimal design yields a minimax-optimal error probability of $\exp\left(-Θ\left(T / H_{G}\right)\right)$, where $H_{G}$ scales proportionally with the dimension $d$. However, this notion of complexity is overly pessimistic, as it is derived from a lower bound in which the arm set consists only of the standard basis vectors, thus masking any potential advantages arising from arm sets with richer geometric structure. To address this, we establish an arm-set-dependent lower bound that, in contrast, holds for any arm set. Motivated by the ideas underlying our lower bound, we propose the Adjacent-optimal design, a specialization of the well-known $\mathcal{X}\mathcal{Y}$-optimal design, and develop the $\textsf{Adjacent-BAI}$ algorithm. We prove that the error probability of $\textsf{Adjacent-BAI}$ matches our lower bound up to constants, verifying the tightness of our lower bound, and establishing the arm-set-dependent complexity of this setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。