新算法能智能利用离线数据,提升强化学习效率。
Offline-Online Reinforcement Learning for Linear Mixture MDPs
- 自适应融合离线数据,根据数据质量决定是否使用
- 在数据有效时,性能优于纯在线学习,且有理论保证
- 适合存在环境变化的现实场景,尤其适合数据有限的领域
我们研究在环境迁移条件下线性混合马尔可夫决策过程(Linear Mixture MDPs)中的离线-在线强化学习。离线阶段数据由未知行为策略采集,可能来自不匹配环境;在线阶段学习者与目标环境交互。提出一种自适应利用离线数据的算法:当离线数据信息量充足(如覆盖充分或环境偏移小),算法可显著优于纯在线学习;当离线数据无用时,安全忽略并保持在线学习性能。建立了显式刻画离线数据价值的后悔上界,并给出近乎匹配的下界。数值实验验证了理论结果。
原文摘要 · Abstract (English)
We study offline-online reinforcement learning in linear mixture Markov decision processes (MDPs) under environment shift. In the offline phase, data are collected by an unknown behavior policy and may come from a mismatched environment, while in the online phase the learner interacts with the target environment. We propose an algorithm that adaptively leverages offline data. When the offline data are informative, either due to sufficient coverage or small environment shift, the algorithm provably improves over purely online learning. When the offline data are uninformative, it safely ignores them and matches the online-only performance. We establish regret upper bounds that explicitly characterize when offline data are beneficial, together with nearly matching lower bounds. Numerical experiments further corroborate our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。