arXiv:2505.18591cs.LGstat.ML2025-05被引 1

用拉普拉斯近似提升元强化学习的不确定性建模能力

Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks

  • 用拉普拉斯近似将元强化学习中的点估计扩展为完整分布
  • 发现传统方法高估置信度且不具一致性,而新方法参数更少
  • 适合需要可靠不确定性估计的机器人、个性化推荐等场景

元强化学习通过在任务分布上训练单一智能体,实现对新任务的快速泛化。从贝叶斯视角看,这相当于对训练任务后验分布进行变分推断。现有方法通常用循环神经网络对任务分布做点估计。本文提出在学习前、中或后使用拉普拉斯近似,将点估计扩展为完整分布,无需修改基础模型架构。借助该近似,可估计非贝叶斯智能体的分布统计量(如熵),发现点估计方法产生过度自信的估计且不满足一致性。与全分布学习相比,本方法性能相当,但参数量显著更少。

原文摘要 · Abstract (English)

Meta-reinforcement learning trains a single reinforcement learning agent on a distribution of tasks to quickly generalize to new tasks outside of the training set at test time. From a Bayesian perspective, one can interpret this as performing amortized variational inference on the posterior distribution over training tasks. Among the various meta-reinforcement learning approaches, a common method is to represent this distribution with a point-estimate using a recurrent neural network. We show how one can augment this point estimate to give full distributions through the Laplace approximation, either at the start of, during, or after learning, without modifying the base model architecture. With our approximation, we are able to estimate distribution statistics (e.g., the entropy) of non-Bayesian agents and observe that point-estimate based methods produce overconfident estimators while not satisfying consistency. Furthermore, when comparing our approach to full-distribution based learning of the task posterior, our method performs on par with variational baselines while having much fewer parameters.

元学习贝叶斯深度学习不确定性估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。