在离线到在线强化学习中,提出新方法实现高效值函数适配。
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
- 基于预训练值函数,设计可证明更优的在线适配算法
- 理论证明在特定条件下样本效率优于纯在线强化学习
- 适用于有高质量离线预训练模型的场景
我们研究在通用函数逼近下从离线到在线强化学习中的值函数适配问题。从一个不完美的离线预训练 $Q$-函数出发,学习者仅通过有限的在线交互适应目标环境。首先,我们通过建立极小极大下界刻画了该设置的难度,表明即使预训练 $Q$-函数接近最优 $Q^/star$,在某些困难实例上,在线适配的效率也无法超过纯在线强化学习。正面来看,我们在一个关于离线预训练值函数的新结构条件下,提出了 O2O-LSVI 算法,其具有问题相关的样本复杂度,可证明优于纯在线强化学习。最后,我们通过神经网络实验验证了该方法的实际有效性。
原文摘要 · Abstract (English)
We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained $Q$-function, the learner aims to adapt it to the target environment using only a limited amount of online interaction. We first characterize the difficulty of this setting by establishing a minimax lower bound, showing that even when the pretrained $Q$-function is close to optimal $Q^\star$, online adaptation can be no more efficient than pure online RL on certain hard instances. On the positive side, under a novel structural condition on the offline-pretrained value functions, we propose O2O-LSVI, an adaptation algorithm with problem-dependent sample complexity that provably improves over pure online RL. Finally, we complement our theory with neural-network experiments that demonstrate the practical effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。