用时间距离衡量状态新颖性,提升稀疏奖励下的探索效率
Episodic Novelty Through Temporal Distance
- 以时间距离为状态相似性度量,避免传统方法在大状态空间失效
- 在多个基准任务上超越现有最佳方法,显著提升探索效果
- 适合解决多轮差异环境中的稀疏奖励问题,如复杂决策场景
在稀疏奖励的上下文马尔可夫决策过程(CMDPs)中,探索仍是重大挑战。现有基于周期的内在动机方法主要依赖计数类方法,在大状态空间中表现不佳,或采用基于相似性的方法,但缺乏合适的状态比较度量。为此,我们提出一种新方法——通过时间距离计算的周期新颖性(ETD),将时间距离作为状态相似性的稳健度量,并结合对比学习准确估计该距离,进而根据当前周期内状态的新颖性生成内在奖励。大量实验表明,ETD在多个基准任务上显著优于现有最优方法,验证了其在稀疏奖励CMDPs中增强探索的有效性。
原文摘要 · Abstract (English)
Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approaches, which are ineffective in large state spaces, or on similarity-based methods that lack appropriate metrics for state comparison. To address these shortcomings, we propose Episodic Novelty Through Temporal Distance (ETD), a novel approach that introduces temporal distance as a robust metric for state similarity and intrinsic reward computation. By employing contrastive learning, ETD accurately estimates temporal distances and derives intrinsic rewards based on the novelty of states within the current episode. Extensive experiments on various benchmark tasks demonstrate that ETD significantly outperforms state-of-the-art methods, highlighting its effectiveness in enhancing exploration in sparse reward CMDPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。