改进自预测强化学习中的自监督损失,提升数据效率。
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
- 在SPR框架中引入终端状态掩码和优先回放加权等策略。
- 在Atari 100k上性能显著提升,且影响SR-SPR与BBF框架。
- 适合关注自监督学习与强化学习融合的研究者。
本研究探究了在SPR框架中对自监督学习(SSL)目标进行特定修改的影响,重点关注终端状态掩码和优先回放加权等未在原始设计中明确考虑的调整。尽管这些修改针对强化学习(RL)特性,但并非适用于所有RL算法。我们评估了六种SPR变体在Atari 100k基准上的表现,包含含与不含上述修改的版本,并测试这些目标在无此类调整的DeepMind Control Suite上的性能。结果表明,在SPR中引入特定的SSL修改能显著提升性能,且该效果可扩展至后续框架如SR-SPR和BBF,凸显了选择合适的SSL目标及其适配策略在实现自预测强化学习数据高效性中的关键作用。
原文摘要 · Abstract (English)
In this study, we investigate the effect of SSL objective modifications within the SPR framework, focusing on specific adjustments such as terminal state masking and prioritized replay weighting, which were not explicitly addressed in the original design. While these modifications are specific to RL, they are not universally applicable across all RL algorithms. Therefore, we aim to assess their impact on performance and explore other SSL objectives that do not accommodate these adjustments like Barlow Twins and VICReg. We evaluate six SPR variants on the Atari 100k benchmark, including versions both with and without these modifications. Additionally, we test the performance of these objectives on the DeepMind Control Suite, where such modifications are absent. Our findings reveal that incorporating specific SSL modifications within SPR significantly enhances performance, and this influence extends to subsequent frameworks like SR-SPR and BBF, highlighting the critical importance of SSL objective selection and related adaptations in achieving data efficiency in self-predictive reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。