探索奖励与记忆架构协同影响强化学习表现,关键在奖励结构设计。
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
- 通过控制实验揭示探索奖励与神经记忆的三类交互模式。
- 奖励结构决定探索与记忆是否互补,影响策略收敛到最优或次优状态。
- 适用于研究强化学习中记忆机制与奖励设计的学者,尤其关注部分可观测环境。
在部分可观测强化学习中,智能体面临双重瓶颈:需探索以发现奖励状态,并将经验存入记忆以优化策略。传统上探索奖励与记忆架构独立评估,其相互作用未被量化;且稀疏奖励的常见定义混淆了信号密度与实际监督内容。本文在三个不同环境中,交叉对比多种神经记忆架构与周期性探索奖励,探究记忆内容获取方式的影响。相同奖励信号导致三种不同交互模式:在需主动发现并无监督保留记忆内容时,奖励放大架构能力差异;当记忆内容为单一受奖励监督的线索时,各架构性能趋于一致;在观察流完全由调度决定时,奖励效果消失。受控奖励操控验证这些模式取决于奖励结构而非密度:密集奖励仅在直接监督所需潜在记忆时才中和探索奖励;轻微可避免的探索惩罚(不改变最优解)会导致策略收敛至次优稳态,而探索奖励可解决此问题。进而提出基于观测锚定的奖励机器形式化奖励稀疏性,区分结构稀疏性(无需任务历史即可重现回报)与潜在稀疏性(单步奖励误定价局部探索动作);该术语体系按任务对记忆的负担组织出三种范式。结果表明探索与记忆是互补关系:奖励促进暴露,记忆将暴露转化为回报。
原文摘要 · Abstract (English)
In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content of memory is acquired. An identical bonus signal yields three distinct interaction patterns: it amplifies architectural capacity differences where memory content must be actively discovered and retained unsupervised; equalizes architectures to a shared ceiling where the content, once sought out, is a single reward-supervised cue; and is null where the observation stream is purely scheduled. Controlled reward manipulations verify that these patterns track reward structure rather than density: a dense reward neutralizes a bonus only if it directly supervises the required latent memory, and a small avoidable penalty on exploratory actions (leaving the optimum unchanged) induces policy convergence to suboptimal stationary states, which either bonus resolves. We then formalize reward sparsity with observation-anchored reward machines, separating structural sparsity (an automaton reproduces the return without the task-required history) from potential sparsity (the one-step reward misprices local exploratory actions); the resulting vocabulary organizes the three regimes by the retention burden each task exposes. Together, these results show exploration and memory are complements, not substitutes: a bonus induces exposure, and only memory converts exposure into return.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。