用量子算法加速连续强化学习中的最优策略探索,提升全局寻优能力。
QuantFPFlow: Quantum Amplitude Estimation for Fokker--Planck Policy Optimisation in Continuous Reinforcement Learning

- 将量子振幅估计算法融入福克-普朗克框架,实现采样效率的平方级提升。
- 在多峰奖励场景中,比SAC更频繁发现全局最优解(33.9% vs 30.7%)。
- 适合需要稳定探索与高维优化的连续控制任务研究者。
我们提出QuantFPFlow,一种将量子振幅估计融入福克-普朗克(FP)形式化随机策略优化的强化学习框架。经典连续空间强化学习需以代价$ ext{O}(1/\varepsilon^2)$估计FP配分函数$Z = int e^{-V(oldsymbol{x})/D}doldsymbol{x}$;QuantFPFlow改用格罗弗放大振幅估计器,达到$ ext{O}(1/\varepsilon)$——实现可证明的二次加速。尽管完整量子加速依赖容错硬件,但此处展示的量子启发式经典模拟已体现$ ext{O}(1/\varepsilon)$算法结构。基于估计的稳态分布$ hostar$生成理论支撑的探索奖励$ Aug = Env + α ext{log}(1/ hostar(s))$,引导智能体向多峰奖励景观中的全局最优区域前进,同时通过FP扩散匹配约束策略方差。在一个专门暴露局部最优失败的连续控制任务上,QuantFPFlow均值奖励为$1{,}295.7 \pm 423.2$,优于软演员-评论家(SAC)的$1{,}284.0 \pm 474.0$,且更频繁发现全局最优(33.9%对30.7%)。训练全程策略熵维持在$H(π)≈6.5$纳特左右,而SAC降至$1.5$纳特,证实扩散匹配有效防止过早收敛。维度实验表明,QuantFPFlow计算复杂度为$ ext{O}(d^{0.35})$,远优于经典方法的$ ext{O}(d^{0.76})$。
原文摘要 · Abstract (English)
We introduce \textbf{QuantFPFlow}, a reinforcement learning framework that integrates quantum amplitude estimation into the Fokker--Planck~(FP) formulation of stochastic policy optimisation. Classical continuous-space RL agents must estimate the FP partition function $Z = \int e^{-V(\mathbf{x})/D}\,d\mathbf{x}$ at cost $\calO(1/\varepsilon^{2})$; QuantFPFlow replaces this with a Grover-amplified amplitude estimator achieving $\calO(1/\varepsilon)$ -- a provable quadratic speedup. While the full quantum acceleration requires fault-tolerant hardware, the quantum-inspired classical simulation demonstrated here already exhibits the $\calO(1/\varepsilon)$ algorithmic structure. The estimated stationary distribution $\rhostar$ drives a theoretically grounded exploration bonus $\Raug = \Renv + α\log(1/\rhostar(s))$. This bonus steers the agent toward globally optimal regions of multimodal reward landscapes while simultaneously constraining policy variance through FP diffusion matching. On a continuous-control task specifically designed to expose local-optima failure, QuantFPFlow achieves mean reward $1{,}295.7 \pm 423.2$ versus $1{,}284.0 \pm 474.0$ for Soft Actor-Critic~(SAC), while discovering the global optimum \textbf{10.4\,\% more frequently} (33.9\,\% vs.\ 30.7\,\%). Policy entropy remains near $H(π)\approx 6.5$\,nats throughout training, whereas SAC collapses to $1.5$\,nats, confirming that FP diffusion matching actively prevents premature convergence. Dimensionality experiments further show computational scaling of $\calO(d^{0.35})$ for QuantFPFlow versus $\calO(d^{0.76})$ for classical FP estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。