arXiv:2409.04374cs.LG2024-09被引 1

用高斯混合模型优化强化学习的Q函数,无需经验数据即超越顶尖方法。

Gaussian-Mixture-Model Q-Functions for Reinforcement Learning by Riemannian Optimization

  • 将高斯混合模型作为Q函数损失的函数逼近器,构建GMM-QF新架构。
  • 在无经验数据条件下,性能超越使用经验数据的深度Q网络等先进方法。
  • 引入黎曼优化框架,为强化学习提供新的理论与计算工具。

本文提出一种新颖的高斯混合模型(GMM)应用方式:将其作为强化学习中Q函数损失的函数逼近器,而非传统概率密度估计。这种新设计称为GMM-QF,被嵌入贝尔曼残差中,构成标准策略迭代中的新型策略评估步骤,实现黎曼优化目标。论文展示了高斯核的均值和协方差矩阵如何从数据中学习,从而开启强化学习对黎曼优化工具箱的利用。数值实验表明,在不使用经验数据的情况下,该方法在基准强化学习任务上优于现有最先进方法,甚至超越依赖经验数据的深度Q网络。

原文摘要 · Abstract (English)

This paper establishes a novel role for Gaussian-mixture models (GMMs) as functional approximators of Q-function losses in reinforcement learning (RL). Unlike the existing RL literature, where GMMs play their typical role as estimates of probability density functions, GMMs approximate here Q-function losses. The new Q-function approximators, coined GMM-QFs, are incorporated in Bellman residuals to promote a Riemannian-optimization task as a novel policy-evaluation step in standard policy-iteration schemes. The paper demonstrates how the hyperparameters (means and covariance matrices) of the Gaussian kernels are learned from the data, opening thus the door of RL to the powerful toolbox of Riemannian optimization. Numerical tests show that with no use of experienced data, the proposed design outperforms state-of-the-art methods, even deep Q-networks which use experienced data, on benchmark RL tasks.

强化学习高斯混合黎曼优化无经验学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。