arXiv:2411.04832cs.AIcs.LG2024-11综述被引 22

揭示深度强化学习中模型适应力下降的根本原因及应对策略

Plasticity Loss in Deep Reinforcement Learning: A Survey

  • 提出统一的可塑性损失定义,系统梳理其成因与病理表现
  • 归纳超50种缓解方法,发现通用正则化效果优于特定领域干预
  • 指出当前评估体系缺陷,建议未来聚焦机制机理研究

可塑性指网络适应变化数据分布的能力,对深度强化学习智能体的成功训练至关重要。可塑性损失导致性能停滞,引发缩放失败、高估偏差和探索不足。为深入理解该问题,本文提出统一定义,分析其驱动因素与病理特征,并将超过50种缓解策略首次系统归类为完整分类体系。分析表明当前评估存在盲区,通用正则化技术往往优于领域专用干预。未来研究应优先关注可塑性损失的内在机制。

原文摘要 · Abstract (English)

Plasticity refers to a network's ability to adapt to changing data distributions, which is crucial for the successful training of deep reinforcement learning agents. Loss of plasticity causes performance plateaus and contributes to scaling failures, overestimation bias, and insufficient exploration. To deepen the understanding of plasticity loss, we propose a unified definition, examine its drivers and pathologies, and organize over 50 mitigation strategies into the first comprehensive taxonomy of the field. Our analysis shows gaps in current evaluation practices and reveals that general regularization techniques often outperform domain-specific interventions. Future research should prioritize understanding the mechanisms underlying plasticity loss.

强化学习可塑性正则化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。