arXiv:2502.00802cs.LGcs.AI2025-02被引 2

用费舍尔信息矩阵缓解强化学习早期内存固化问题

Fisher-Guided Selective Forgetting: Mitigating The Primacy Bias in Deep Reinforcement Learning

  • 基于费舍尔信息矩阵识别学习中的记忆与重组阶段
  • 提出选择性遗忘机制,使早期经验不主导模型更新
  • 在复杂任务中显著提升性能,适合强化学习研究者

深度强化学习系统常因过拟合早期经验而产生首要偏差(Primacy Bias, PB),严重影响学习效率和最终表现,尤其在复杂环境中。本文从费舍尔信息矩阵(FIM)视角系统分析PB,发现学习过程中FIM迹呈现特定模式,揭示关键的记忆固化与参数重组织阶段。基于此,提出费舍尔引导的选择性遗忘(FGSF)方法,利用参数空间的几何结构动态调整网络权重,防止早期经验占据主导。在DeepMind Control Suite(DMC)环境上的实验表明,FGSF在复杂任务中持续优于基线。研究还分析了PB对策略网络与价值网络的不同影响,验证了回放比率对偏差的加剧作用,并评估了简单噪声注入的有效性。结果深化了对PB的理解,提供了基于FIM的几何视角与实用缓解策略,推动深度强化学习发展。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) systems often tend to overfit to early experiences, a phenomenon known as the primacy bias (PB). This bias can severely hinder learning efficiency and final performance, particularly in complex environments. This paper presents a comprehensive investigation of PB through the lens of the Fisher Information Matrix (FIM). We develop a framework characterizing PB through distinct patterns in the FIM trace, identifying critical memorization and reorganization phases during learning. Building on this understanding, we propose Fisher-Guided Selective Forgetting (FGSF), a novel method that leverages the geometric structure of the parameter space to selectively modify network weights, preventing early experiences from dominating the learning process. Empirical results across DeepMind Control Suite (DMC) environments show that FGSF consistently outperforms baselines, particularly in complex tasks. We analyze the different impacts of PB on actor and critic networks, the role of replay ratios in exacerbating the effect, and the effectiveness of even simple noise injection methods. Our findings provide a deeper understanding of PB and practical mitigation strategies, offering a FIM-based geometric perspective for advancing DRL.

强化学习偏差缓解参数优化FIM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。