arXiv:2603.03523cs.LGmath.OC2026-03被引 2

提出Q测度学习,高效实现连续状态强化学习。

Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence

  • 用核积分重建值函数,通过测度学习避免存储无限维估计
  • 每轮迭代仅需O(n)内存与计算,支持在线学习
  • 理论保证收敛性,适合高维连续控制任务

我们研究无限时域折扣马尔可夫决策过程中的强化学习,其中数据来自单一轨迹的马尔可夫行为策略。为避免维护无穷维函数估计,提出新型Q测度学习:学习在访问过的状态-动作对上支撑的带符号经验测度,并通过核积分重构动作价值估计。该方法联合估计行为链的平稳分布与Q测度,采用耦合随机逼近,实现基于权重的高效实现,每轮迭代内存与计算复杂度均为O(n)。在行为链均匀遍历条件下,证明了诱导出的Q函数几乎必然以一致范数收敛到核平滑贝尔曼算子的不动点。同时,给出了该极限与最优Q*之间的近似误差界,其依赖于核带宽。通过在两商品库存控制场景中的强化学习实验评估算法性能。

原文摘要 · Abstract (English)

We study reinforcement learning in infinite-horizon discounted Markov decision processes with continuous state spaces, where data are generated online from a single trajectory under a Markovian behavior policy. To avoid maintaining an infinite-dimensional, function-valued estimate, we propose the novel Q-Measure-Learning, which learns a signed empirical measure supported on visited state-action pairs and reconstructs an action-value estimate via kernel integration. The method jointly estimates the stationary distribution of the behavior chain and the Q-measure through coupled stochastic approximation, leading to an efficient weight-based implementation with $O(n)$ memory and $O(n)$ computation cost per iteration. Under uniform ergodicity of the behavior chain, we prove almost sure sup-norm convergence of the induced Q-function to the fixed point of a kernel-smoothed Bellman operator. We also bound the approximation error between this limit and the optimal $Q^*$ as a function of the kernel bandwidth. To assess the performance of our proposed algorithm, we conduct RL experiments in a two-item inventory control setting.

强化学习连续状态测度学习核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。