arXiv:2506.14479stat.MLcs.LG2025-06被引 1

改进了贝叶斯强化学习中的采样策略,提升推荐系统决策效率

Adaptive Data Augmentation for Thompson Sampling

  • 设计自适应样本增强机制,动态优化参数估计
  • 在多个数据集上实现更优累积回报,显著降低损失
  • 无需假设上下文分布,适合真实场景应用

在线性上下文老虎机问题中,目标是选择能最大化累计奖励的动作,其奖励被建模为未知参数的线性函数。尽管汤普森采样在实践中表现良好,但未能达到最优后悔界。本文通过提出一种新型估计器,结合为高效参数学习而设计的假设样本的自适应增强与耦合,实现了接近极小极大最优的汤普森采样。该估计器无需依赖上下文分布假设,即可准确预测所有臂的奖励。实验结果表明,所提方法具有稳健性能,并显著优于现有方法。

原文摘要 · Abstract (English)

In linear contextual bandits, the objective is to select actions that maximize cumulative rewards, modeled as a linear function with unknown parameters. Although Thompson Sampling performs well empirically, it does not achieve optimal regret bounds. This paper proposes a nearly minimax optimal Thompson Sampling for linear contextual bandits by developing a novel estimator with the adaptive augmentation and coupling of the hypothetical samples that are designed for efficient parameter learning. The proposed estimator accurately predicts rewards for all arms without relying on assumptions for the context distribution. Empirical results show robust performance and significant improvement over existing methods.

强化学习贝叶斯优化在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。