arXiv:2502.12227cs.LGcs.AI2025-02

针对已知支持的多项分布奖励,提出更优的最优臂识别方法。

Identifying the Best Transition Law

  • 基于经验似然和独立偏差界,结合联合概率向量估计。
  • 在不同结构复杂度场景中,性能优于传统非参数方法。
  • 适合关注多臂赌博机中高效探索策略的研究者。

受马尔可夫决策过程递归学习启发,本文研究了每个臂的奖励来自已知支撑的多项分布时的最优臂识别问题。比较了包括LUCB在内的多种策略的性能,其中一种利用该知识进行概率估计。在第一种情况下,采用经典的非参数置信区间;在第二种情况下,先对各维度独立使用霍夫丁和伯恩斯坦偏差界,再用经验似然方法(EL-LUCB)处理联合概率向量。通过在不同结构复杂度场景下的模拟实验,验证了这些方法的有效性。

原文摘要 · Abstract (English)

Motivated by recursive learning in Markov Decision Processes, this paper studies best-arm identification in bandit problems where each arm's reward is drawn from a multinomial distribution with a known support. We compare the performance { reached by strategies including notably LUCB without and with use of this knowledge. } In the first case, we use classical non-parametric approaches for the confidence intervals. In the second case, where a probability distribution is to be estimated, we first use classical deviation bounds (Hoeffding and Bernstein) on each dimension independently, and then the Empirical Likelihood method (EL-LUCB) on the joint probability vector. The effectiveness of these methods is demonstrated through simulations on scenarios with varying levels of structural complexity.

多臂赌博机最优臂识别经验似然置信区间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。