arXiv:2412.08031cs.LG2024-12被引 2

在分组强化学习中,找出满足条件且收益最高的选项。

Constrained Best Arm Identification in Grouped Bandits

  • 将每个臂拆分为多个属性,仅当所有属性均达标时才视为可行
  • 提出基于置信区间的近优算法,在固定置信度下高效识别最优可行臂
  • 适合需要严格筛选条件的推荐系统或医疗决策场景

我们研究一种分组多臂老虎机设置,其中每个臂由多个独立的子臂(即属性)组成,每个属性具有独立的随机回报。我们设定约束:只有当某个臂的所有属性的均值回报均超过指定阈值时,该臂才被视为可行。目标是在固定置信度下,从所有可行臂中识别出属性平均回报最高的臂。我们首先刻画了任意策略性能的理论极限。随后,提出一种基于置信区间的近优策略,并提供该策略的理论保证。通过模拟对比,验证了所提策略与两种经适当修改的动作消除法相比的优越性。

原文摘要 · Abstract (English)

We study a grouped bandit setting where each arm comprises multiple independent sub-arms referred to as attributes. Each attribute of each arm has an independent stochastic reward. We impose the constraint that for an arm to be deemed feasible, the mean reward of all its attributes should exceed a specified threshold. The goal is to find the arm with the highest mean reward averaged across attributes among the set of feasible arms in the fixed confidence setting. We first characterize a fundamental limit on the performance of any policy. Following this, we propose a near-optimal confidence interval-based policy to solve this problem and provide analytical guarantees for the policy. We compare the performance of the proposed policy with that of two suitably modified versions of action elimination via simulations.

强化学习多臂老虎机约束优化决策算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。