arXiv:2506.10664stat.MLcs.LG2025-06

提出新算法解决策略持续迭代中的离线学习难题。

Sequential Off-Policy Learning with Logarithmic Smoothing

  • 结合对数平滑与在线PAC-Bayesian理论设计新算法
  • 在策略持续更新时收敛更快,性能显著优于旧方法
  • 适合需要反复更新策略的实际系统应用

离线学习允许从历史交互数据中训练策略。以往工作多聚焦于批量设置,即从单一行为策略生成的数据中学习。然而在实际系统中,策略会反复更新并重新部署,每次利用所有已收集数据训练,并生成未来更新的新交互。这种序列式离线学习场景虽常见,但理论研究仍不充分。本文提出并分析一种简单算法,将对数平滑(LS)估计与在线PAC-Bayesian工具结合。我们进一步证明,对LS进行合理调整可在弱条件下提升性能并加速收敛。所提算法具有广泛适应性:在批量情形下达到当前最优水平,在序列更新时显著超越现有方法。实验验证了序列框架的优势及算法的有效性。

原文摘要 · Abstract (English)

Off-policy learning enables training policies from logged interaction data. Most prior work considers the batch setting, where a policy is learned from data generated by a single behavior policy. In real systems, however, policies are updated and redeployed repeatedly, each time training on all previously collected data while generating new interactions for future updates. This sequential off-policy learning setting is common in practice but remains largely unexplored theoretically. In this work, we present and study a simple algorithm for sequential off-policy learning, combining Logarithmic Smoothing (LS) estimation with online PAC-Bayesian tools. We further show that a principled adjustment to LS improves performance and accelerates convergence under mild conditions. The algorithms introduced generalise previous work: they match state-of-the-art offline approaches in the batch case and substantially outperform them when policies are updated sequentially. Empirical evaluations highlight both the benefits of the sequential framework and the strength of the proposed algorithms.

离线学习策略迭代对数平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。