arXiv:2502.03061cs.LG2025-02被引 3

带后动作上下文的最优动作识别,提升决策效率

Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context

  • 引入后动作上下文信息优化多臂赌博机中的最优动作识别
  • 分离与非分离场景下均实现最优样本复杂度
  • 适合需要高效探索的强化学习与推荐系统应用

我们提出在随机多臂赌博机环境中,基于固定置信度设置的最佳动作识别(BAI)新问题,即在每次执行动作后,学习者不仅获得奖励,还获得一个后动作上下文。该上下文提供额外信息,显著促进决策过程。我们分析两种类型的后动作上下文:(i) 分离型,其中奖励仅依赖于上下文;(ii) 非分离型,其中奖励同时依赖于动作和上下文。针对两种情形,我们推导出实例相关的样本复杂度下界,并提出渐近最优的算法。在分离型设置中,提出一种名为G-tracking的新采样规则,利用上下文空间的几何结构直接追踪上下文而非动作;在非分离型设置中,证明Track-and-Stop算法可扩展至该场景。理论上并实证表明,忽略后动作上下文的算法为次优。实验结果展示所提方法优于现有最先进方法。

原文摘要 · Abstract (English)

We introduce the problem of best arm identification (BAI) with post-action context, a new BAI problem in a stochastic multi-armed bandit environment and the fixed-confidence setting. The problem addresses the scenarios in which the learner receives a post-action context in addition to the reward after playing each action. This post-action context provides additional information that can significantly facilitate the decision process. We analyze two different types of the post-action context: (i) separator, where the reward depends solely on the context, and (ii) non-separator, where the reward depends on both the action and the context. For both cases, we derive instance-dependent lower bounds on the sample complexity and propose algorithms that asymptotically achieve the optimal sample complexity. For the separator setting, we propose a novel sampling rule called G-tracking, which uses the geometry of the context space to directly track the contexts rather than the actions. For the non-separator setting, we do so by demonstrating that the Track-and-Stop algorithm can be extended to this setting. Moreover, in both settings, we theoretically and empirically show that algorithms that ignore the post-action context are sub-optimal. Finally, our empirical results showcase the advantage of our approaches compared to the state of the art.

强化学习多臂赌博机探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。