arXiv:2509.19098cs.LGmath.ST2025-09

利用先验样本提升多臂赌博机学习效率,显著降低累积损失

Asymptotically Optimal Problem-Dependent Bandit Policies for Transfer Learning

  • 基于源分布样本设计迁移策略,通过距离约束优化目标选择
  • 在高斯场景下实现理论最优的渐近累积后悔率
  • 适合有可靠先验数据的强化学习与推荐系统应用

我们研究非上下文多臂赌博机在迁移学习设置下的问题:在任何动作执行前,学习者会获得每个源分布 ν'_k 的 N'_k 个独立同分布样本,而真实目标分布 ν_k 与源分布之间的已知距离满足 d_k(ν_k, ν'_k) ≤ L_k。在此框架下,我们首先推导出一个依赖具体问题的渐近下界,扩展了经典的 Lai-Robbins 结果,纳入了迁移参数 (d_k, L_k, N'_k)。随后提出 KL-UCB-Transfer 策略,该简单索引策略在高斯情形下可达到此新下界。最后通过模拟验证,当源与目标分布足够接近时,KL-UCB-Transfer 显著优于无先验基线。

原文摘要 · Abstract (English)

We study the non-contextual multi-armed bandit problem in a transfer learning setting: before any pulls, the learner is given N'_k i.i.d. samples from each source distribution nu'_k, and the true target distributions nu_k lie within a known distance bound d_k(nu_k, nu'_k) <= L_k. In this framework, we first derive a problem-dependent asymptotic lower bound on cumulative regret that extends the classical Lai-Robbins result to incorporate the transfer parameters (d_k, L_k, N'_k). We then propose KL-UCB-Transfer, a simple index policy that matches this new bound in the Gaussian case. Finally, we validate our approach via simulations, showing that KL-UCB-Transfer significantly outperforms the no-prior baseline when source and target distributions are sufficiently close.

多臂赌博机迁移学习强化学习统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。