提出新方法让混合策略在连续动作强化学习中更稳定高效
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

- 用边缘化重参数化估计器解决混合策略方差高的问题
- 在多个基准环境上性能超越传统混合策略,接近高斯策略表现
- 为复杂策略设计提供实用方案,适合追求鲁棒性的强化学习研究者
混合策略在连续动作强化学习中理论上比单模态策略更具灵活性,但其实际优势尚不明确。当前多数先进算法未采用混合策略,引发根本疑问:这种表示开销是否值得?我们证明更高灵活性可提升解的质量与熵鲁棒性。然而标准算法如SAC未能利用这些优势,核心原因是混合策略缺乏低方差的重参数化技巧(高斯策略享有此优势)。为此,我们提出边缘化重参数化(MRP)估计器,证明其方差低于标准似然比(LR)方法。在Gym MuJoCo、DeepMind Control Suite和MetaWorld上的实验表明,采用MRP的混合策略显著优于LR方法,并达到甚至超过高斯策略的表现。此外,我们观察到若干情况下MRP混合策略展现出明显优势。本工作厘清了其中权衡关系,使MRP混合策略从理论构想变为实用工具。
原文摘要 · Abstract (English)
Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain elusive. Mixture policies are notably absent from most state-of-the-art algorithms, raising a fundamental question: Is the added representational overhead useful? We show that increased flexibility can theoretically enhance solution quality and entropy robustness. Yet standard algorithms like SAC do not leverage these advantages. A core issue is the lack of a low-variance reparameterization trick for mixtures, a luxury Gaussian policies enjoy. We propose a marginalized reparameterization (MRP) estimator to address this, proving it offers lower variance than the standard likelihood-ratio (LR) approach. Our experiments across Gym MuJoCo, DeepMind Control Suite, and MetaWorld show that MRP mixture policies significantly outperform their LR ones, and reach parity (sometimes better) with Gaussian counterparts. In addition, we do find several cases where MRP mixture policies exhibit clear empirical advantages. In this paper, we provide a clearer understanding of the trade-offs involved, elevating MRP mixture policies from theoretical curiosity to a practical tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。