arXiv:2512.20974cs.LGcs.AI2025-12

用可学习基函数的广义线性模型提升深度贝叶斯强化学习的泛化能力

Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions

  • 引入可学习基函数的广义线性模型,实现任务参数的精确贝叶斯推断
  • 在MuJoCo和MetaWorld上性能提升最高达1.8倍,超越现有元强化学习方法
  • 首次建立任务表示距离与核对应关系的闭式联系,适合研究高效贝叶斯强化学习者

贝叶斯强化学习(BRL)作为元强化学习(Meta-RL)的一个子类,通过显式引入贝叶斯任务参数到转移和奖励模型中,提供了具有理论基础的泛化框架。然而,传统BRL方法假设转移和奖励模型的形式已知。尽管近期深度BRL方法引入了模型学习以解决此问题,但直接将神经网络应用于联合数据与任务参数仍需变分推断,常导致任务表示模糊,影响最终策略性能。为此,我们提出基于可学习基函数的深度贝叶斯强化学习广义线性模型(GLiBRL)。该方法实现了任务参数与模型噪声的完全可解析贝叶斯推断,并支持精确的边缘似然评估以学习转移和奖励模型。其置换不变的精确贝叶斯推断结构可无缝集成于在线与离线RL算法。进一步证明,GLiBRL在任务表示的$/mathcal{L}_2$距离与任务样本间的核对应之间存在闭式关系,据我们所知这是首个针对在线深度BRL的结构性结果。在代表性元强化学习方法对比中,GLiBRL在MuJoCo和MetaWorld基准上性能提升最高达1.8×。

原文摘要 · Abstract (English)

Bayesian Reinforcement Learning (BRL), a subclass of Meta-Reinforcement Learning (Meta-RL), provides a principled framework for generalisation by explicitly incorporating Bayesian task parameters into transition and reward models. However, classical BRL methods assume known forms of transition and reward models. While recent deep BRL methods incorporate model learning to address this, applying neural networks directly to joint data and task parameters necessitates variational inference. This often yields indistinct task representations, compromising the resulting BRL policies. To overcome these limitations, we introduce Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions (GLiBRL). Our approach features fully tractable Bayesian inference over task parameters and model noise, alongside exact marginal likelihood evaluation for learning transition and reward models. The permutation-invariant nature of exact Bayesian inference in GLiBRL enables seamless integration with both on-policy and off-policy RL algorithms. We further show that GLiBRL admits a closed-form relationship between the $\mathcal{L}_2$ distance of its task representations and empirical kernel-based correspondence between task samples, which is to our knowledge the first such structural result for online deep BRL. GLiBRL is compared against representative and recent Meta-RL methods, and improves state-of-the-art performance on both MuJoCo and MetaWorld benchmarks by up to 1.8$\times$.

贝叶斯强化学习元强化学习广义线性模型可学习基函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。