arXiv:2410.14665cs.LGcs.AI2024-10被引 1

利用预存数据提升在线强化学习性能,理论证明近最优。

Online Reinforcement Learning with Passive Memory

  • 结合预收集数据与在线学习,优化决策过程。
  • 理论证明后悔值接近最小最大最优,性能显著提升。
  • 适用于连续与离散状态空间,适合有历史数据的场景。

本文研究一种利用环境预收集数据(被动记忆)的在线强化学习算法。通过引入被动记忆,算法在在线交互中表现更优,并给出了近最小最大最优的后悔率理论保证。结果表明,被动记忆的质量决定了后悔值的次优程度。该方法和理论分析适用于连续与离散状态-动作空间。

原文摘要 · Abstract (English)

This paper considers an online reinforcement learning algorithm that leverages pre-collected data (passive memory) from the environment for online interaction. We show that using passive memory improves performance and further provide theoretical guarantees for regret that turns out to be near-minimax optimal. Results show that the quality of passive memory determines sub-optimality of the incurred regret. The proposed approach and results hold in both continuous and discrete state-action spaces.

强化学习在线学习记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。