预训练视觉表征在模型基于强化学习中效果不佳,反而不如从零训练。
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
- 在基于模型的强化学习中测试多种预训练视觉表征
- 预训练表征未提升样本效率,且在分布外场景泛化更差
- 数据多样性与网络结构对泛化性能影响最大
视觉强化学习(RL)通常需要大量数据。与无模型RL相比,基于模型的强化学习(MBRL)可通过规划实现高效的数据利用。此外,RL在真实任务中缺乏泛化能力。已有研究表明,引入预训练视觉表征(PVRs)可提升样本效率和泛化性能。尽管PVRs在无模型RL中被广泛研究,其在MBRL中的潜力仍基本未被探索。本文在挑战性控制任务中对一系列PVRs进行基准测试,评估其在基于模型的代理中的数据效率、泛化能力及不同表征特性的影响。结果表明,令人意外的是,当前的PVRs在MBRL中并未比从零学习表示更具样本效率,且在分布外(OOD)设置下泛化能力更差。通过分析训练的动力学模型质量,我们解释了这一现象。此外,我们发现数据多样性和网络架构是影响OOD泛化性能的最关键因素。
原文摘要 · Abstract (English)
Visual Reinforcement Learning (RL) methods often require extensive amounts of data. As opposed to model-free RL, model-based RL (MBRL) offers a potential solution with efficient data utilization through planning. Additionally, RL lacks generalization capabilities for real-world tasks. Prior work has shown that incorporating pre-trained visual representations (PVRs) enhances sample efficiency and generalization. While PVRs have been extensively studied in the context of model-free RL, their potential in MBRL remains largely unexplored. In this paper, we benchmark a set of PVRs on challenging control tasks in a model-based RL setting. We investigate the data efficiency, generalization capabilities, and the impact of different properties of PVRs on the performance of model-based agents. Our results, perhaps surprisingly, reveal that for MBRL current PVRs are not more sample efficient than learning representations from scratch, and that they do not generalize better to out-of-distribution (OOD) settings. To explain this, we analyze the quality of the trained dynamics model. Furthermore, we show that data diversity and network architecture are the most important contributors to OOD generalization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。