arXiv:2605.23138quant-phcs.AI2026-05

用强化学习生成量子态初始值,加速变分量子算法收敛。

Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning

论文配图:Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning
图 1 · 摘自论文原文
  • 将门选择建模为序列决策问题,用Transformer驱动的搜索策略选最优克利福德门
  • 22量子比特、1370参数下平均能量精度提升3.17倍,最佳结果提升45倍
  • 适合需要快速初始化的变分量子算法研究者,尤其在复杂问题上表现突出

变分量子算法(VQAs)有望实现实用量子优势,但其优化受限于平底区和众多局部极小值。尽管可经典模拟的克利福德电路可用于预热VQAs以加速收敛,现有基于启发式的方法在庞大组合搜索空间中难以扩展。为此,我们提出CRiSP(用于状态制备的克利福德强化学习代理),将离散前缀选择建模为序列决策问题。CRiSP利用基于Transformer的策略通过自对弈训练,结合神经引导蒙特卡洛树搜索,在固定参数旋转前插入学习到的克利福德门。这使得高质量初始态的构建完全可通过多项式时间的经典稳定器模拟实现,且不改变底层电路结构。通过渐进扩展搜索范围的课程学习策略,该代理能高效扩展至深层电路。在最多22量子比特、1370个参数的QAOA基准测试中,CRiSP相比现有最优克利福德初始化方法,平均能量精度提升3.17倍(最高达45.02倍),最佳能量精度提升2.44倍(最高达16.01倍)。在VQE任务上的评估进一步验证了该框架的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren plateaus and numerous local minima. While classically simulable Clifford circuits can warm-start VQAs to accelerate convergence, existing heuristic-based initialization methods struggle to scale within vast combinatorial search spaces. To overcome this bottleneck, we propose CRiSP (a Clifford Reinforcement Learning agent for State Preparation), a framework that formulates discrete prefix selection as a sequential decision-making problem. CRiSP utilizes Neural-Guided Monte Carlo Tree Search, driven by a Transformer-based policy trained via self-play, to insert learned Clifford gates before fixed parameterized rotations. This enables the construction of high-quality initial states entirely through polynomial-time classical stabilizer simulation without altering the underlying circuit architecture. By integrating a curriculum learning strategy that progressively expands the search horizon, the agent efficiently scales to deep circuits. Evaluated on QAOA benchmarks of up to $22$ qubits and $1{,}370$ parameters, CRiSP outperforms state-of-the-art Clifford initialization methods by a mean of $3.17\times$ (max $45.02\times$) in average energy accuracy and $2.44\times$ (max $16.01\times$) in best-achieved energy accuracy. Assessments on VQE tasks further demonstrate the framework's robustness and generalizability.

量子计算强化学习变分算法状态制备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。