arXiv:2505.20017stat.MLcs.LG2025-05

放宽噪声独立假设,提出新算法应对依赖噪声的线性强化学习问题。

Linear Bandits with Non-i.i.d. Noise

  • 基于信心序列与乐观原则设计新算法
  • 理论证明后悔上界与依赖衰减速率相关
  • 适合研究非独立噪声场景的强化学习者

我们研究线性随机博弈问题,放松了观测噪声独立同分布的标准假设。作为替代,允许噪声在轮次间为亚高斯但存在依赖,且依赖随时间衰减。为此,我们利用近期提出的序贯概率分配归约方法,构建新的置信序列,并基于乐观面对不确定性原则设计带宽算法。我们给出了该算法的后悔上界,其形式取决于观测间依赖强度的衰减速率。在其他结果中,我们证明当观测噪声几何混合时,上界可恢复标准速率,仅差一个混合时间因子。

原文摘要 · Abstract (English)

We study the linear stochastic bandit problem, relaxing the standard i.i.d. assumption on the observation noise. As an alternative to this restrictive assumption, we allow the noise terms across rounds to be sub-Gaussian but interdependent, with dependencies that decay over time. To address this setting, we develop new confidence sequences using a recently introduced reduction scheme to sequential probability assignment, and use these to derive a bandit algorithm based on the principle of optimism in the face of uncertainty. We provide regret bounds for the resulting algorithm, expressed in terms of the decay rate of the strength of dependence between observations. Among other results, we show that our bounds recover the standard rates up to a factor of the mixing time for geometrically mixing observation noise.

强化学习线性带宽非独立噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。