用时序相关潜空间提升强化学习探索效率,更抗噪声与随机性。
A Temporally Correlated Latent Exploration for Reinforcement Learning
- 基于动作条件潜空间和时序相关性设计新内在奖励机制
- 在Minigrid和随机Atari环境中显著提升探索鲁棒性
- 适合需要高效探索的复杂随机环境任务
高效探索仍是深度强化学习中的长期难题。现有方法依赖环境外在奖励,或引入内在奖励增强探索,但易受噪声电视(Noisy TV)和随机性影响。为此,本文提出时序相关潜空间探索(TeCLE),一种新颖的内在奖励设计,通过动作条件潜空间与时序相关性建模状态概率分布,避免对不可预测状态分配过高内在奖励,有效应对上述问题。与以往在动作选择中注入时序相关性不同,本方法将时序相关性用于内在奖励计算,实验表明其注入的时序相关性直接决定智能体的探索行为。在基准环境如Minigrid和随机Atari上的测试验证了TeCLE对噪声和随机性的鲁棒性。据我们所知,这是首个将动作条件潜空间与时序相关性结合用于好奇心驱动探索的方法。
原文摘要 · Abstract (English)
Efficient exploration remains one of the longstanding problems of deep reinforcement learning. Instead of depending solely on extrinsic rewards from the environments, existing methods use intrinsic rewards to enhance exploration. However, we demonstrate that these methods are vulnerable to Noisy TV and stochasticity. To tackle this problem, we propose Temporally Correlated Latent Exploration (TeCLE), which is a novel intrinsic reward formulation that employs an action-conditioned latent space and temporal correlation. The action-conditioned latent space estimates the probability distribution of states, thereby avoiding the assignment of excessive intrinsic rewards to unpredictable states and effectively addressing both problems. Whereas previous works inject temporal correlation for action selection, the proposed method injects it for intrinsic reward computation. We find that the injected temporal correlation determines the exploratory behaviors of agents. Various experiments show that the environment where the agent performs well depends on the amount of temporal correlation. To the best of our knowledge, the proposed TeCLE is the first approach to consider the action conditioned latent space and temporal correlation for curiosity-driven exploration. We prove that the proposed TeCLE can be robust to the Noisy TV and stochasticity in benchmark environments, including Minigrid and Stochastic Atari.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。