arXiv:2605.00488cs.LG2026-05被引 37

平衡探索与收益,提升多臂赌博机的决策效率

Trading off rewards and errors in multi-armed bandits

论文配图:Trading off rewards and errors in multi-armed bandits
图 1 · 摘自论文原文
  • 设计算法在探索与收益间动态权衡
  • 理论证明可实现双目标的最优折中
  • 适合需兼顾学习精度与收益场景

在多臂赌博机问题中,最常被探索的臂信息量最大,而奖励最大化通常只选择最优臂。本文研究准确识别臂均值与累积奖励之间的权衡,提出一种具有后悔界保证的算法,可在两个目标间进行插值。我们同时给出了上界和下界,并通过实验验证了其有效性。

原文摘要 · Abstract (English)

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm means accurately and accumulating reward, and present an algorithm with regret guarantees that interpolates between the two objectives. We provide both upper and lower bounds and validate empirically.

强化学习多臂赌博机优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。