arXiv:2501.12627cs.LG2025-01被引 2

用混合奖励机制提升强化学习的探索效率与多样性

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

  • 设计可灵活融合多种内在奖励的HIRE框架
  • 在多个基准上显著提升复杂环境下的探索效率
  • 适合研究探索算法或应对稀疏奖励问题的开发者

内在奖励塑造已成为解决强化学习中困难探索与稀疏奖励环境的主流方法。尽管单一内在奖励(如基于好奇或新颖性的方法)已证明有效,但往往限制了探索的多样性和效率。同时,多种内在奖励结合的潜力与原理仍不充分。为此,我们提出HIRE(Hybrid Intrinsic REward)——一种通过有意识融合策略构建混合内在奖励的灵活且优雅的框架。借助HIRE,我们在多个基准上系统分析了混合内在奖励在通用及无监督强化学习中的应用。大量实验表明,HIRE能显著提升复杂动态环境中的探索效率、多样性以及技能习得能力。

原文摘要 · Abstract (English)

Intrinsic reward shaping has emerged as a prevalent approach to solving hard-exploration and sparse-rewards environments in reinforcement learning (RL). While single intrinsic rewards, such as curiosity-driven or novelty-based methods, have shown effectiveness, they often limit the diversity and efficiency of exploration. Moreover, the potential and principle of combining multiple intrinsic rewards remains insufficiently explored. To address this gap, we introduce HIRE (Hybrid Intrinsic REward), a flexible and elegant framework for creating hybrid intrinsic rewards through deliberate fusion strategies. With HIRE, we conduct a systematic analysis of the application of hybrid intrinsic rewards in both general and unsupervised RL across multiple benchmarks. Extensive experiments demonstrate that HIRE can significantly enhance exploration efficiency and diversity, as well as skill acquisition in complex and dynamic settings.

强化学习探索策略奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。