arXiv:2607.20010cs.LGcs.CE2026-07

用概率推断框架提升强化学习的不确定性估计能力

Generalized Kalman filter based temporal difference reinforcement learning

论文配图:Generalized Kalman filter based temporal difference reinforcement learning
图 1 · 摘自论文原文
  • 将价值函数视为随机变量,基于条件期望构建非线性非高斯强化学习框架
  • 同时估计价值函数均值与方差,实现全程不确定性量化
  • 适用于复杂系统建模,适合关注可靠性与鲁棒性的研究者

本文提出一种基于条件期望理论的广义时序差分(TD)强化学习框架。将价值函数和动作价值函数视为不确定量,其估计被形式化为随机推理问题。不同于传统基于卡尔曼滤波的TD方法依赖线性高斯假设,本框架直接从条件期望出发,自然拓展至非线性模型和非高斯分布。所提方法递归估计价值函数的条件期望及其二阶概率矩,从而在学习过程中持续量化其不确定性。为获得可计算算法,采用多项式混沌展开或集合近似对随机问题进行离散化,实现对底层随机变量的有效表示。在两个最优控制问题上验证:线性质量-弹簧-阻尼系统和封闭腔体中的非线性热传导问题。数值结果表明该方法能准确估计价值函数及其不确定性,将经典卡尔曼基时序差分学习推广至更广泛的随机系统。

原文摘要 · Abstract (English)

In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochastic inference problem. Unlike classical Kalman-based temporal-difference learning, which relies on linear-Gaussian assumptions, the proposed formulation is derived directly from the conditional expectation framework and naturally extends to nonlinear models and non-Gaussian probability distributions. The proposed method recursively estimates not only the conditional expectation of the value function but also its second probabilistic moment, thereby quantifying the uncertainty associated with the learned value function throughout the learning process. To obtain a computationally tractable algorithm, the stochastic problem is discretized using either polynomial chaos expansions or ensemble-based approximations, providing efficient representations of the underlying random variables. The proposed framework is demonstrated on two optimal control problems: a linear mass--spring--damper system and a nonlinear heat conduction problem in a closed cavity. The numerical examples illustrate the capability of the proposed method to accurately estimate both the value function and its associated uncertainty, while extending classical Kalman-based temporal-difference learning to a broader class of stochastic systems.

强化学习不确定性概率推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。