arXiv:2502.17666cs.LGcs.AI2025-02被引 8

用Q-learning提升离线上下文强化学习,性能平均提高30%

Yes, Q-learning Helps Offline In-Context RL

  • 在离线上下文RL中直接优化强化学习目标
  • 相比主流方法算法蒸馏,平均性能提升30%
  • 适合关注离线强化学习与上下文学习融合的研究者

现有离线上下文强化学习(ICRL)方法主要依赖监督训练目标,但在离线强化学习场景中存在局限。本文探索在离线ICRL框架中引入强化学习目标。在150多个由GridWorld和MuJoCo环境生成的数据集上实验表明,直接优化强化学习目标可使性能平均比广泛采用的算法蒸馏(Algorithm Distillation, AD)提升约30%,且在不同数据覆盖度、结构、专家水平和环境复杂性下均有效。在更具挑战性的XLand-MiniGrid环境中,强化学习目标使性能达到AD的两倍。此外,在价值学习中加入保守性策略在几乎所有测试场景中进一步提升了性能。结果表明,将ICRL的学习目标与强化学习的奖励最大化目标对齐至关重要,证实离线强化学习是推动ICRL发展的有前途方向。

原文摘要 · Abstract (English)

Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.

强化学习离线学习上下文学习Q-learning

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。