对比量子与经典算法在控制任务中的表现,发现经典方法更优但量子模型潜力更大。
Hybrid Quantum-Classical Policy Gradient for Adaptive Control of Cyber-Physical Systems: A Comparative Study of VQC vs. MLP
- 用多层感知机和变分量子电路分别训练控制策略,对比性能差异。
- 经典模型平均收益达498.7,量子模型仅14.6,差距显著且受硬件限制。
- 量子模型参数少、训练时间略增,适合未来低资源量子设备部署。
本研究对比了经典与量子强化学习范式在基准控制环境中的表现,评估其收敛性、观测噪声下的鲁棒性及计算效率。采用多层感知机(MLP)作为经典基线,变分量子电路(VQC)作为量子对比模型,均在CartPole-v1环境中训练500轮。实验表明,经典MLP达到近最优策略,平均回报为498.7 ± 3.2,训练过程稳定;而VQC学习能力有限,平均回报仅为14.6 ± 4.8,主要受限于电路深度与量子比特连通性。噪声鲁棒性分析显示,MLP在高斯扰动下性能衰减平缓,而VQC在相同噪声水平下敏感度更高。尽管最终性能较低,但VQC参数量显著减少,训练时间略有增加,展现出在低资源量子处理器上的可扩展潜力。结果表明,当前控制基准中经典神经策略仍占主导,但一旦硬件噪声与表达力瓶颈被克服,量子增强架构可能带来效率优势。
原文摘要 · Abstract (English)
The comparative evaluation between classical and quantum reinforcement learning (QRL) paradigms was conducted to investigate their convergence behavior, robustness under observational noise, and computational efficiency in a benchmark control environment. The study employed a multilayer perceptron (MLP) agent as a classical baseline and a parameterized variational quantum circuit (VQC) as a quantum counterpart, both trained on the CartPole-v1 environment over 500 episodes. Empirical results demonstrated that the classical MLP achieved near-optimal policy convergence with a mean return of 498.7 +/- 3.2, maintaining stable equilibrium throughout training. In contrast, the VQC exhibited limited learning capability, with an average return of 14.6 +/- 4.8, primarily constrained by circuit depth and qubit connectivity. Noise robustness analysis further revealed that the MLP policy deteriorated gracefully under Gaussian perturbations, while the VQC displayed higher sensitivity at equivalent noise levels. Despite the lower asymptotic performance, the VQC exhibited significantly lower parameter count and marginally increased training time, highlighting its potential scalability for low-resource quantum processors. The results suggest that while classical neural policies remain dominant in current control benchmarks, quantum-enhanced architectures could offer promising efficiency advantages once hardware noise and expressivity limitations are mitigated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。