arXiv:2505.07911cs.LGcs.AI2025-05综述被引 1

融合贝叶斯与强化学习,提升智能体决策的数据效率与可解释性

Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review

  • 用贝叶斯方法量化不确定性,增强决策可靠性
  • 系统梳理五类融合策略,覆盖模型内外多种场景
  • 适合关注安全、高效智能决策的研究者参考

贝叶斯推断在智能体(如机器人/仿真代理)决策中相较于传统数据驱动黑箱神经网络具有数据高效、泛化能力强、可解释性高和安全性好等优势,这些优势直接或间接源于其对不确定性的量化能力。然而,目前缺乏全面综述系统总结贝叶斯推断在强化学习中的进展。本文聚焦于贝叶斯与强化学习的结合,涵盖五大内容:1)适用于决策的贝叶斯方法(如贝叶斯规则、共轭模型、变分推断、贝叶斯优化、贝叶斯深度学习、主动学习、生成模型、元学习与持续学习);2)经典融合方式(基于模型的强化学习、免模型强化学习及逆强化学习);3)最新潜在方法组合;4)从数据效率、泛化能力、可解释性与安全性角度对方法进行分析比较;5)深入探讨六类复杂强化学习问题(未知奖励、部分可观测、多智能体、多任务、非线性非高斯、层次化),并总结贝叶斯方法在数据收集、处理与策略学习各阶段的作用,为构建更优决策策略提供路径。

原文摘要 · Abstract (English)

Bayesian inference has many advantages in decision making of agents (e.g. robotics/simulative agent) over a regular data-driven black-box neural network: Data-efficiency, generalization, interpretability, and safety where these advantages benefit directly/indirectly from the uncertainty quantification of Bayesian inference. However, there are few comprehensive reviews to summarize the progress of Bayesian inference on reinforcement learning (RL) for decision making to give researchers a systematic understanding. This paper focuses on combining Bayesian inference with RL that nowadays is an important approach in agent decision making. To be exact, this paper discusses the following five topics: 1) Bayesian methods that have potential for agent decision making. First basic Bayesian methods and models (Bayesian rule, Bayesian learning, and Bayesian conjugate models) are discussed followed by variational inference, Bayesian optimization, Bayesian deep learning, Bayesian active learning, Bayesian generative models, Bayesian meta-learning, and lifelong Bayesian learning. 2) Classical combinations of Bayesian methods with model-based RL (with approximation methods), model-free RL, and inverse RL. 3) Latest combinations of potential Bayesian methods with RL. 4) Analytical comparisons of methods that combine Bayesian methods with RL with respect to data-efficiency, generalization, interpretability, and safety. 5) In-depth discussions in six complex problem variants of RL, including unknown reward, partial-observability, multi-agent, multi-task, non-linear non-Gaussian, and hierarchical RL problems and the summary of how Bayesian methods work in the data collection, data processing and policy learning stages of RL to pave the way for better agent decision-making strategies.

贝叶斯推断强化学习智能体决策可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。