用扩散模型提升离线到在线强化学习的探索效率与可靠性
Efficient and Uncertainty-Aware Diffusion Framework for Offline-to-Online Reinforcement Learning

- 离线阶段用扩散模型蒸馏快速采样的策略和状态转移模型
- 在线阶段通过拉普拉斯近似量化不确定性,平衡探索与利用
- 在多个环境中显著提升在线回报,优于现有基线方法
离线到在线强化学习(O2O-RL)利用预训练的离线策略以减少昂贵的在线交互。尽管数据高效,但其易受离线与在线分布差异的影响。现有方法通过在扩散模型采样轨迹上微调策略来缓解这一问题。受此启发,我们提出DUAL:一种高效的扩散不确定性感知框架用于离线到在线强化学习。DUAL在离线阶段利用扩散模型的先验知识,蒸馏出可快速采样的扩散动作策略和状态转移模型。在在线阶段,DUAL采用拉普拉斯近似与距离状态偏移检测,通过不确定性量化优化探索与利用的权衡。我们形式化证明,基于拉普拉斯近似的动作损失可作为认知不确定性的一致估计代理。实验表明,DUAL在多个环境设置下均显著提升在线期望回报,优于O2O-RL基线。
原文摘要 · Abstract (English)
Offline-to-Online Reinforcement Learning (O2O-RL) leverages an offline, pre-trained policy to minimize costly online interactions. Although data-efficient, O2O-RL is susceptible to shifts between offline and online distributions. Existing work aims to mitigate the harm of this shift by finetuning the policy on trajectory data sampled from a diffusion model. Inspired by this line of work, we propose DUAL: an efficient \textbf{D}iffusion \textbf{U}ncertainty-\textbf{A}ware framework for offline-to-online reinforcement \textbf{L}earning. DUAL utilizes the prior knowledge of the diffusion model to distill a fast-sampling diffusion actor policy and transition model in the offline phase. DUAL also employs a Laplace approximation and distance transition-state-shift detection, thereby using uncertainty quantification to improve exploration versus exploitation in the online phase. We formally show that our actor loss with the Laplace approximation provides a proxy for a principled estimate of epistemic uncertainty. Empirically, DUAL improves the online expected return over O2O-RL baselines across multiple settings and environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。