主动采样上下文可显著减少学习所需样本数。
Active Learning for Stochastic Contextual Linear Bandits

- 通过主动选择上下文进行采样,提升学习效率。
- 理论证明可比最优最坏情况快√d倍,d为维度。
- 在华法林剂量预测等任务中显著降低样本需求。
随机上下文线性老虎机的核心目标是高效学习近似最优策略。以往算法通过策略性采样动作来学习策略,但对上下文的采样方式是被动的,直接从上下文分布中随机抽取。然而,在在线内容推荐、调查研究和临床试验等实际场景中,从业者可根据对上下文分布的先验知识主动采样或招募上下文。尽管存在主动学习的潜力,当前对随机上下文线性老虎机中战略性上下文采样的研究仍不足。本文提出一种新算法,通过策略性采样上下文-动作对的奖励来学习近似最优策略。我们给出了实例相关理论保证,证明所提主动上下文采样策略可比最小最大率快至√d倍,其中d为线性维度。实验表明,该算法在华法林剂量预测和笑话推荐等任务中显著减少了学习近似最优策略所需的样本数量。
原文摘要 · Abstract (English)
A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions but naively (passively) sampling contexts from the underlying context distribution. However, in many practical scenarios -- including online content recommendation, survey research, and clinical trials -- practitioners can actively sample or recruit contexts based on prior knowledge of the context distribution. Despite this potential for active learning, the role of strategic context sampling in stochastic contextual linear bandits is underexplored. We propose an algorithm that learns a near-optimal policy by strategically sampling rewards of context-action pairs. We prove instance-dependent theoretical guarantees demonstrating that our active context sampling strategy can improve over the minimax rate by up to a factor of $\sqrt{d}$, where $d$ is the linear dimension. We show empirically that our algorithm reduces the number of samples needed to learn a near-optimal policy, in tasks such as warfarin dose prediction and joke recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。