基于结构信息设计新探索机制,提升强化学习采样效率。
Effective Exploration Based on the Structural Information Principles
- 用结构互信息捕捉状态动作间动态相关性,构建编码树。
- 通过最大化条件结构熵,实现状态动作空间更优覆盖。
- 在多个基准上性能超基线,样本效率提升达60.25%。
传统信息论为强化学习中的表征学习和熵最大化探索提供了基础,但现有方法多聚焦于随机变量的不确定性建模,忽视了状态与动作空间的内在结构。本文提出基于结构信息原理的有效探索框架SI2E:定义变量间的结构互信息以突破单变量局限,并提出新型嵌入原则,捕获与动态相关的状态-动作表示。SI2E分析策略在状态-动作对上的价值差异,最小化结构熵以构建分层状态-动作结构(即编码树)。在此结构上,定义并最大化价值条件下的结构熵,设计内在奖励机制,避免冗余转移,增强状态-动作空间覆盖。理论证明了SI2E与经典信息论方法的关联,验证其合理性与优势。在MiniGrid、MetaWorld及DeepMind Control Suite等基准上的综合评估表明,SI2E在最终性能和样本效率上显著优于当前最先进探索基线,最大提升分别达37.63%和60.25%。
原文摘要 · Abstract (English)
Traditional information theory provides a valuable foundation for Reinforcement Learning, particularly through representation learning and entropy maximization for agent exploration. However, existing methods primarily concentrate on modeling the uncertainty associated with RL's random variables, neglecting the inherent structure within the state and action spaces. In this paper, we propose a novel Structural Information principles-based Effective Exploration framework, namely SI2E. Structural mutual information between two variables is defined to address the single-variable limitation in structural information, and an innovative embedding principle is presented to capture dynamics-relevant state-action representations. The SI2E analyzes value differences in the agent's policy between state-action pairs and minimizes structural entropy to derive the hierarchical state-action structure, referred to as the encoding tree. Under this tree structure, value-conditional structural entropy is defined and maximized to design an intrinsic reward mechanism that avoids redundant transitions and promotes enhanced coverage in the state-action space. Theoretical connections are established between SI2E and classical information-theoretic methodologies, highlighting our framework's rationality and advantage. Comprehensive evaluations in the MiniGrid, MetaWorld, and DeepMind Control Suite benchmarks demonstrate that SI2E significantly outperforms state-of-the-art exploration baselines regarding final performance and sample efficiency, with maximum improvements of 37.63% and 60.25%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。