用熵信息引导智能体突破状态边界,提升稀疏奖励下的探索效率
Explore Beyond the Boundary Using Entropic Information

- 基于熵信息识别状态分布边界并赋予内在奖励
- 在稀疏延迟奖励环境中显著提升探索性能
- 适合需要高效探索的强化学习场景
在强化学习中,稀疏且延迟的奖励使得探索面临重大挑战,因为缺乏足够的反馈来指导学习过程。为解决这一问题,需要在状态空间中进行广泛探索以发现有价值的奖励信号。本文提出一种名为熵信息探索(ENTINEX)的新方法,通过激励智能体探索状态分布边界来增强探索能力。该方法利用熵信息有效识别边界,并为这些边界分配内在奖励。大量实验表明,ENTINEX在具有稀疏和延迟奖励的环境中持续提升探索表现,优于现有探索方法,充分证明其在复杂奖励场景下的有效性。
原文摘要 · Abstract (English)
In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process. Addressing this issue requires extensive exploration in the state space to discover valuable reward signals. In this paper, we propose Entropic Information for Exploration (ENTINEX), a novel method that enhances exploration by incentivizing agents to explore beyond the boundaries of the state distribution. ENTINEX achieves this by assigning intrinsic rewards to these boundaries, leveraging entropic information to identify them effectively. Through extensive experimentation, we demonstrate that ENTINEX consistently improves exploration performance in environments characterized by sparse and delayed rewards. Our experimental results show that ENTINEX outperforms existing exploration methods, highlighting its effectiveness in both sparse and delayed reward scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。