用强化学习优化量子态制备,提升参数化电路生成效率
Reinforcement Learning for Parameterized Quantum State Preparation: A Comparative Study
- 将强化学习从离散门选择扩展到连续旋转角优化
- 单阶段PPO在2-10量子比特上成功率83%-99%,λ≤4时有效
- 推荐单阶段PPO方案,兼顾性能与计算效率
本文将定向量子电路合成(DQCS)拓展至参数化量子态制备,引入连续单量子比特旋转(Rx, Ry, Rz)。对比两种训练方式:单阶段代理同时选择门类型、作用量子比特和旋转角度;双阶段方法先生成离散电路,再用Adam优化角度。基于Gymnasium与PennyLane,评估PPO与A2C在2-10量子比特系统上的表现,目标复杂度参数λ为1-5。结果显示,A2C无法学习有效策略,而稳定超参数下的PPO成功实现(单阶段:学习率约5×10⁻⁴,自保真度误差阈值0.01;双阶段:学习率约10⁻⁴)。两类方法均可可靠重构计算基态(成功率83%-99%)与贝尔态(61%-77%),但当λ≈3-4时可扩展性饱和,十量子比特目标即使λ=2也无法实现。双阶段方法仅带来微弱精度提升,但运行时间增加约三倍。因此,在固定算力下推荐使用单阶段PPO策略,并提供具体合成电路,与经典变分基线对比,指出未来可提升可扩展性的方向。
原文摘要 · Abstract (English)
We extend directed quantum circuit synthesis (DQCS) with reinforcement learning from purely discrete gate selection to parameterized quantum state preparation with continuous single-qubit rotations \(R_x\), \(R_y\), and \(R_z\). We compare two training regimes: a one-stage agent that jointly selects the gate type, the affected qubit(s), and the rotation angle; and a two-stage variant that first proposes a discrete circuit and subsequently optimizes the rotation angles with Adam using parameter-shift gradients. Using Gymnasium and PennyLane, we evaluate Proximal Policy Optimization (PPO) and Advantage Actor--Critic (A2C) on systems comprising two to ten qubits and on targets of increasing complexity with \(λ\) ranging from one to five. Whereas A2C does not learn effective policies in this setting, PPO succeeds under stable hyperparameters (one-stage: learning rate approximately \(5\times10^{-4}\) with a self-fidelity-error threshold of 0.01; two-stage: learning rate approximately \(10^{-4}\)). Both approaches reliably reconstruct computational basis states (between 83\% and 99\% success) and Bell states (between 61\% and 77\% success). However, scalability saturates for \(λ\) of approximately three to four and does not extend to ten-qubit targets even at \(λ=2\). The two-stage method offers only marginal accuracy gains while requiring around three times the runtime. For practicality under a fixed compute budget, we therefore recommend the one-stage PPO policy, provide explicit synthesized circuits, and contrast with a classical variational baseline to outline avenues for improved scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。