arXiv:2509.14585cs.LGmath.OC2025-09被引 2

用稀疏高斯混合模型构建可解释的在线强化学习框架

Online reinforcement learning via sparse Gaussian mixture model Q-functions

  • 基于稀疏高斯混合模型设计可解释的在线策略迭代方法
  • 参数量少于深度强化学习,且在低参数下仍保持强性能
  • 适合对模型可解释性与高效性有要求的研究者

本文提出一种基于稀疏高斯混合模型Q函数(S-GMM-QFs)的结构化、可解释的在线强化学习策略迭代框架。相较于早期离线训练的GMM-QFs,该框架利用流式数据促进探索,通过哈达玛超参数化实现稀疏化以控制模型复杂度,在避免过拟合的同时保持表达能力。S-GMM-QFs的参数空间天然具备黎曼流形结构,支持在平滑目标上进行在线梯度下降更新。数值实验表明,S-GMM-QFs在标准基准上表现媲美甚至超越密集型深度强化学习方法,同时使用显著更少的参数;在低参数情形下,其性能优于稀疏化的深度强化学习方法。

原文摘要 · Abstract (English)

This paper introduces a structured and interpretable online policy-iteration framework for reinforcement learning (RL), built around the novel class of sparse Gaussian mixture model Q-functions (S-GMM-QFs). Extending earlier work that trained GMM-QFs offline, the proposed framework develops an online scheme that leverages streaming data to encourage exploration. Model complexity is regulated through sparsification by Hadamard overparametrization, which mitigates overfitting while preserving expressiveness. The parameter space of S-GMM-QFs is naturally endowed with a Riemannian manifold structure, allowing for principled parameter updates via online gradient descent on a smooth objective. Numerical experiments show that S-GMM-QFs match or even outperform dense deep RL (DeepRL) methods on standard benchmarks while using significantly fewer parameters. Moreover, they maintain strong performance even in low-parameter regimes where sparsified DeepRL methods fail to generalize.

强化学习稀疏模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。