用稀疏高斯混合模型提升强化学习的效率与可解释性
Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning
- 通过哈达玛超参数化构建稀疏高斯混合Q函数,实现几何可解释的在线优化
- 参数量少于深度强化学习方法,且每步更新后性能提升更快
- 适合需要高效、可解释模型的实时强化学习场景
本文提出一种基于稀疏高斯混合模型Q函数(S-GMM-QFs)的在线、离策略策略迭代框架。该框架通过哈达玛超参数化实现稀疏化,利用平滑正则化促进几何可解释性,并在黎曼流形上进行梯度下降优化,有效处理非平稳数据与经验分布不匹配问题。模型从大初始池中自适应识别有意义的成分,各组件的均值与协方差参数显式表征其在状态-动作空间中的几何角色。数值实验表明,S-GMM-QFs在参数量显著减少的情况下,性能可媲美甚至超越深度强化学习方法,且在低参数条件下仍保持强泛化能力。
原文摘要 · Abstract (English)
This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure of the parameter space while handling distributional mismatch through experience replay. S-GMM-QFs are introduced via Hadamard overparametrization, enabling interpretable sparsification through smooth regularization that facilitates Riemannian-based optimization. Overparametrization allows the framework to adaptively identify meaningful components from a large initial pool, yielding sparse models where interpretability emerges naturally from geometry: each component's parameters (means and covariances) explicitly encode its geometric role in the ambient state-action space. These geometric roles are learned through online gradient descent on a smooth objective over a (Cartesian-product) Riemannian manifold. Numerical tests demonstrate that S-GMM-QFs match or exceed deep RL methods while using substantially fewer parameters and achieving faster improvement per observed transition. Notably, parameter efficiency and interpretability combine to maintain strong generalization in low-parameter regimes where sparsified deep RL approaches degrade.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。