用各向同性高斯表示提升强化学习训练稳定性
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
- 采用各向同性高斯正则化,使表征分布更稳定
- 在多种任务中减少表征坍塌与训练不稳,提升性能
- 适合应对目标变化快、数据分布漂移的强化学习场景
深度强化学习系统常因非平稳性导致训练不稳定,即学习目标和数据分布随时间演变。我们证明,在非平稳目标下,各向同性高斯嵌入具有理论优势:能稳定追踪时变目标、在固定方差预算下实现最大熵,并促进所有表征维度的均衡使用,从而增强智能体的适应性与稳定性。基于此,我们提出草图式各向同性高斯正则化(Sketched Isotropic Gaussian Regularization),在训练中引导表征趋向各向同性高斯分布。实验证明,该方法在多种领域均有效提升非平稳条件下的性能,同时减少表征坍塌、神经元失活和训练不稳定性,且计算开销极低。
原文摘要 · Abstract (English)
Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data distributions evolve over time. We show that under non-stationary targets, isotropic Gaussian embeddings are provably advantageous. In particular, they induce stable tracking of time-varying targets for linear readouts, achieve maximal entropy under a fixed variance budget, and encourage a balanced use of all representational dimensions--all of which enable agents to be more adaptive and stable. Building on this insight, we propose the use of Sketched Isotropic Gaussian Regularization for shaping representations toward an isotropic Gaussian distribution during training. We demonstrate empirically, over a variety of domains, that this simple and computationally inexpensive method improves performance under non-stationarity while reducing representation collapse, neuron dormancy, and training instability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。