arXiv:2511.06895cs.LGcs.AI2025-11被引 1

发现深度强化学习中存在模型容量过大会提升泛化能力的现象。

On The Presence of Double-Descent in Deep Reinforcement Learning

  • 用策略熵衡量策略不确定性,分析不同模型规模下的表现
  • 训练过程中出现明显的双下降曲线,过参数化后性能持续提升
  • 为设计更通用、鲁棒的强化学习智能体提供新思路

深度强化学习(DRL)中过参数化模型的泛化能力在非平稳环境下仍不明确。本文基于无模型的演员-评论家框架,系统研究了不同模型容量下的双下降(DD)现象。通过信息论指标——策略熵(Policy Entropy)来度量训练过程中的策略不确定性。初步结果表明,在训练周期层面出现了清晰的双下降曲线;当策略进入第二下降区时,策略熵显著且持续下降。这种熵的衰减表明,过参数化起到了隐式正则化作用,引导策略走向损失曲面中更稳健、更平坦的极小值。该发现确立了双下降在DRL中的存在性,并提供了基于信息论的智能体设计机制,有助于提升其泛化性、可迁移性和鲁棒性。

原文摘要 · Abstract (English)

The double descent (DD) paradox, where over-parameterized models see generalization improve past the interpolation point, remains largely unexplored in the non-stationary domain of Deep Reinforcement Learning (DRL). We present preliminary evidence that DD exists in model-free DRL, investigating it systematically across varying model capacity using the Actor-Critic framework. We rely on an information-theoretic metric, Policy Entropy, to measure policy uncertainty throughout training. Preliminary results show a clear epoch-wise DD curve; the policy's entrance into the second descent region correlates with a sustained, significant reduction in Policy Entropy. This entropic decay suggests that over-parameterization acts as an implicit regularizer, guiding the policy towards robust, flatter minima in the loss landscape. These findings establish DD as a factor in DRL and provide an information-based mechanism for designing agents that are more general, transferable, and robust.

强化学习双下降策略熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。