通过无监督生成数据提升离线强化学习性能
Unsupervised Data Generation for Offline Reinforcement Learning: A Perspective from Model
- 从模型视角理论证明数据分布差异决定离线强化学习性能差距
- 无监督训练多策略可最小化未知任务的最坏后悔值
- 提出UDG方法在无任务依赖场景下生成更优训练数据
离线强化学习近年受到广泛关注,但其性能受限于分布外问题,该问题在在线强化学习中可通过反馈修正。以往研究多关注限制算法在分布内采样,较少关注批量数据的影响。本文从基于模型的离线强化学习优化视角,首次理论建立批量数据与算法性能之间的联系。在温和假设下,我们证明:行为策略与最优策略生成的状态-动作对分布之间的距离,决定了基于模型的离线强化学习策略与最优策略间的性能差距。其次,在任务无关设置下,一系列由无监督强化学习训练得到的策略,可最小化性能差距中的最坏后悔值。受此启发,我们提出无监督数据生成方法UDG,用于在任务无关设置下生成并选择适合离线训练的数据。实验表明,UDG在解决未知任务时优于监督式数据生成方法。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) recently gains growing interests from RL researchers. However, the performance of offline RL suffers from the out-of-distribution problem, which can be corrected by feedback in online RL. Previous offline RL research focuses on restricting the offline algorithm in in-distribution even in-sample action sampling. In contrast, fewer work pays attention to the influence of the batch data. In this paper, we first build a bridge over the batch data and the performance of offline RL algorithms theoretically, from the perspective of model-based offline RL optimization. We draw a conclusion that, with mild assumptions, the distance between the state-action pair distribution generated by the behavioural policy and the distribution generated by the optimal policy, accounts for the performance gap between the policy learned by model-based offline RL and the optimal policy. Secondly, we reveal that in task-agnostic settings, a series of policies trained by unsupervised RL can minimize the worst-case regret in the performance gap. Inspired by the theoretical conclusions, UDG (Unsupervised Data Generation) is devised to generate data and select proper data for offline training under tasks-agnostic settings. Empirical results demonstrate that UDG can outperform supervised data generation on solving unknown tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。