arXiv:2606.04305cs.LGstat.ML2026-06

利用离线数据提升线性强化学习早期表现,随时间自动减少依赖

Offline-to-Online Learning in Linear Bandits

论文配图:Offline-to-Online Learning in Linear Bandits
图 1 · 摘自论文原文
  • 结合离线数据与在线探索,动态调整策略
  • 在线交互次数增长时,后悔值呈亚线性下降
  • 适合有历史数据但需在线适应的场景

我们研究在随机线性带宽设置下,利用额外离线数据的在线学习问题。尽管该问题在实践中频繁出现,但在结构化环境中的离线到在线权衡仍不明确。我们提出一种线性带宽算法,通过早期依赖离线数据并随时间推移逐步增加探索来平衡这一权衡。理论分析表明,该方法在纯在线和纯离线方案中均表现优异:相对于最优在线动作的后悔值随在线交互次数增长呈亚线性下降;而相对于离线参考的后悔值则随离线样本数量增加而减小。实验结果进一步验证了其在多种问题参数下的有效性。

原文摘要 · Abstract (English)

We study online learning with an additional offline dataset in the stochastic linear bandit setting. Although this problem arises frequently in practice, the offline-to-online tradeoff remains poorly understood in structured environments. We propose a linear bandit algorithm that balances this tradeoff: it relies on offline data during early rounds, and increasingly favors exploration as the horizon grows. We establish regret bounds showing that our method is simultaneously competitive with both purely online and purely offline solutions. In particular, it achieves sublinear regret relative to the optimal action in the number of online interactions, while its regret relative to an offline reference decreases as the number of offline samples grows. Empirical results further demonstrate its effectiveness across various problem parameters.

强化学习线性带宽离线数据在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。