arXiv:2410.07638cs.LGcs.AI2024-10NeurIPS被引 6

提出新模型与算法,高效识别非平稳线性老虎机中的最优动作

Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits

  • 设计并行双子算法,主动检测环境突变并对齐上下文
  • 理论证明样本复杂度近似最小化,且实验证明高效
  • 适合动态环境中需快速定位最优策略的研究者

我们提出一种新颖的分段平稳线性老虎机(PSLB)模型,其中环境在每个突变点随机采样未知分布的上下文,各动作的质量由其在所有上下文上的平均回报衡量。上下文分布和突变点均对智能体未知。我们设计了分段平稳ε-最优动作识别+(PSεBAI⁺)算法,可保证以至少1−δ的概率找到ε-最优动作,并实现最少采样次数。该算法由两个并行子程序构成:主动检测突变点并对齐上下文的PSεBAI,以及基础的NεBAI。理论分析表明,PSεBAI⁺具有有限的期望采样复杂度,且其复杂度仅比下界大对数因子,达到几乎极小极大最优。数值实验对比基线算法,验证了其高效性。分析与实验共同表明,其性能优势源于PSεBAI中精细的突变检测与上下文对齐机制。

原文摘要 · Abstract (English)

We propose a {\em novel} piecewise stationary linear bandit (PSLB) model, where the environment randomly samples a context from an unknown probability distribution at each changepoint, and the quality of an arm is measured by its return averaged over all contexts. The contexts and their distribution, as well as the changepoints are unknown to the agent. We design {\em Piecewise-Stationary $\varepsilon$-Best Arm Identification$^+$} (PS$\varepsilon$BAI$^+$), an algorithm that is guaranteed to identify an $\varepsilon$-optimal arm with probability $\ge 1-δ$ and with a minimal number of samples. PS$\varepsilon$BAI$^+$ consists of two subroutines, PS$\varepsilon$BAI and {\sc Naïve $\varepsilon$-BAI} (N$\varepsilon$BAI), which are executed in parallel. PS$\varepsilon$BAI actively detects changepoints and aligns contexts to facilitate the arm identification process. When PS$\varepsilon$BAI and N$\varepsilon$BAI are utilized judiciously in parallel, PS$\varepsilon$BAI$^+$ is shown to have a finite expected sample complexity. By proving a lower bound, we show the expected sample complexity of PS$\varepsilon$BAI$^+$ is optimal up to a logarithmic factor. We compare PS$\varepsilon$BAI$^+$ to baseline algorithms using numerical experiments which demonstrate its efficiency. Both our analytical and numerical results corroborate that the efficacy of PS$\varepsilon$BAI$^+$ is due to the delicate change detection and context alignment procedures embedded in PS$\varepsilon$BAI.

强化学习在线决策最优动作识别非平稳环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。