arXiv:2410.10132cs.LGstat.ML2024-10ICLR被引 8

提出稳定哈达玛记忆,提升强化学习中长期记忆能力。

Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement Learning

  • 用哈达玛积动态调节记忆,高效删冗余、留关键信息。
  • 在POPGym等长时程任务上超越现有方法,显著提升性能。
  • 适合需要长期依赖和自适应记忆的强化学习场景。

部分可观测环境中的有效决策依赖于稳健的记忆管理。尽管深度学习记忆模型在监督学习中表现优异,但在具有部分可观测性和长期依赖的强化学习环境中仍存在不足:难以高效捕捉相关历史信息、灵活适应变化观测,且长期训练中更新不稳定。本文通过统一框架理论分析现有记忆模型的局限性,提出一种新型强化学习记忆模型——稳定哈达玛记忆(Stable Hadamard Memory)。该模型通过哈达玛积实现计算高效的内存校准与更新,动态清除过时经验,强化关键记忆。在元强化学习、长时程信用分配及POPGym等挑战性基准任务上,本方法显著优于当前最先进记忆模型,展现出对长期演化上下文的优异处理能力。

原文摘要 · Abstract (English)

Effective decision-making in partially observable environments demands robust memory management. Despite their success in supervised learning, current deep-learning memory models struggle in reinforcement learning environments that are partially observable and long-term. They fail to efficiently capture relevant past information, adapt flexibly to changing observations, and maintain stable updates over long episodes. We theoretically analyze the limitations of existing memory models within a unified framework and introduce the Stable Hadamard Memory, a novel memory model for reinforcement learning agents. Our model dynamically adjusts memory by erasing no longer needed experiences and reinforcing crucial ones computationally efficiently. To this end, we leverage the Hadamard product for calibrating and updating memory, specifically designed to enhance memory capacity while mitigating numerical and learning challenges. Our approach significantly outperforms state-of-the-art memory-based methods on challenging partially observable benchmarks, such as meta-reinforcement learning, long-horizon credit assignment, and POPGym, demonstrating superior performance in handling long-term and evolving contexts.

强化学习记忆机制长时依赖哈达玛积

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。