通过自适应Wasserstein正则化提升强化学习的稳定性。
Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning
- 在评论家损失中加入自适应Wasserstein正则项,稳定训练过程。
- 理论证明评论家均方误差收敛速度达O(1/k),优于传统方法。
- 适合追求训练稳定性的深度强化学习研究者使用。
我们提出Wasserstein自适应价值估计(WAVE),一种通过自适应Wasserstein正则化增强深度强化学习稳定性的方法。该方法通过在评论家损失函数中引入可自适应调整的Wasserstein正则项,解决演员-评论家算法固有的不稳定性问题。理论上证明了WAVE在评论家均方误差上达到O(1/k)的收敛速率,并通过基于Wasserstein的正则化提供了稳定性保障。利用Sinkhorn近似实现计算高效性,使正则化强度能根据智能体表现自动调节。理论分析与实验结果表明,相比标准演员-评论家方法,WAVE实现了更优性能。
原文摘要 · Abstract (English)
We present Wasserstein Adaptive Value Estimation for Actor-Critic (WAVE), an approach to enhance stability in deep reinforcement learning through adaptive Wasserstein regularization. Our method addresses the inherent instability of actor-critic algorithms by incorporating an adaptively weighted Wasserstein regularization term into the critic's loss function. We prove that WAVE achieves $\mathcal{O}\left(\frac{1}{k}\right)$ convergence rate for the critic's mean squared error and provide theoretical guarantees for stability through Wasserstein-based regularization. Using the Sinkhorn approximation for computational efficiency, our approach automatically adjusts the regularization based on the agent's performance. Theoretical analysis and experimental results demonstrate that WAVE achieves superior performance compared to standard actor-critic methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。