arXiv:2511.04856cs.LGquant-ph2025-11

用量子模型提升连续动作强化学习效率,避免传统方法的不稳定性。

Quantum Boltzmann Machines for Sample-Efficient Reinforcement Learning

  • 结合经典先验与量子分布,构建混合量子-经典模型
  • 可解析计算连续变量梯度,直接嵌入演员-评论家算法
  • 采样替代全局最大化,解决连续控制中的训练不稳问题

我们提出了理论基础坚实的连续半量子玻尔兹曼机(CSQBMs),支持连续动作强化学习。通过在可见单元上引入指数族先验,并在隐藏单元上采用量子玻尔兹曼分布,CSQBMs 构建了一种混合量子-经典模型,在减少量子比特需求的同时保持强表达能力。关键优势在于,可对连续变量进行解析梯度计算,从而直接集成到演员-评论家算法中。基于此,我们提出一种连续Q-learning框架,以高效采样代替全局最大化,有效克服了连续控制中的不稳定性问题。

原文摘要 · Abstract (English)

We introduce theoretically grounded Continuous Semi-Quantum Boltzmann Machines (CSQBMs) that supports continuous-action reinforcement learning. By combining exponential-family priors over visible units with quantum Boltzmann distributions over hidden units, CSQBMs yield a hybrid quantum-classical model that reduces qubit requirements while retaining strong expressiveness. Crucially, gradients with respect to continuous variables can be computed analytically, enabling direct integration into Actor-Critic algorithms. Building on this, we propose a continuous Q-learning framework that replaces global maximization by efficient sampling from the CSQBM distribution, thereby overcoming instability issues in continuous control.

强化学习量子机器学习连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。