arXiv:2506.13111cs.LGcs.AI2025-06中稿 · IEEE Statistical S…被引 1

用高斯过程扩散模型缓解强化学习过拟合问题

Overcoming Overfitting in Reinforcement Learning via Gaussian Process Diffusion Policy

  • 结合高斯过程与扩散模型构建新策略,提升泛化能力
  • 在行走机器人测试中,分布偏移下性能提升67.74%~123.18%
  • 适合需应对环境变化的强化学习系统研究者

强化学习面临的核心挑战之一是难以适应由不确定性引起的数据分布变化。这一问题在使用深度神经网络作为决策器或策略的RL系统中尤为突出,长期在固定环境中训练后易产生过拟合。为此,本文提出高斯过程扩散策略(GPDP),将扩散模型与高斯过程回归(GPR)结合以表示策略。GPR引导扩散模型生成最大化已学Q函数的动作,类似强化学习中的策略改进。此外,GPR基于核函数的特性提升了测试时分布偏移下的探索效率,增加了发现新行为的可能性,有效缓解过拟合。在Walker2d基准上的仿真结果表明,本方法在分布偏移条件下优于现有先进算法,使强化学习目标函数提升约67.74%至123.18%,且在正常条件下表现相当。

原文摘要 · Abstract (English)

One of the key challenges that Reinforcement Learning (RL) faces is its limited capability to adapt to a change of data distribution caused by uncertainties. This challenge arises especially in RL systems using deep neural networks as decision makers or policies, which are prone to overfitting after prolonged training on fixed environments. To address this challenge, this paper proposes Gaussian Process Diffusion Policy (GPDP), a new algorithm that integrates diffusion models and Gaussian Process Regression (GPR) to represent the policy. GPR guides diffusion models to generate actions that maximize learned Q-function, resembling the policy improvement in RL. Furthermore, the kernel-based nature of GPR enhances the policy's exploration efficiency under distribution shifts at test time, increasing the chance of discovering new behaviors and mitigating overfitting. Simulation results on the Walker2d benchmark show that our approach outperforms state-of-the-art algorithms under distribution shift condition by achieving around 67.74% to 123.18% improvement in the RL's objective function while maintaining comparable performance under normal conditions.

强化学习过拟合扩散模型高斯过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。