用强化学习找量子电路的高效去纠缠测量方案。
Reinforced Disentanglers on Random Unitary Circuits
- 用PPO算法优化量子电路中的投影测量位置
- 所需测量数远低于传统相变理论预测
- 揭示最优测量模式,适合量子纠错研究
我们采用近端策略优化(PPO)算法,在两比特门呈砖墙结构的随机克利福德电路中搜索高效的去纠缠测量方案。去纠缠测量定义为插入连续纠缠层之间的投影测量集合。高效去纠缠测量指在总投影数最少的前提下,使最终态的平均冯诺依曼熵最小。将表示测量配置的二值矩阵作为状态,比特翻转操作作为动作,基于最终态平均熵和测量配置给予奖励,使智能体从无测量初始状态逐步学习到最优测量策略。结果表明,去纠缠所需测量数显著少于测量诱导相变研究中的数值结果。此外,强化学习过程揭示了最优测量的规律性模式,这是以往相变研究无法实现的。
原文摘要 · Abstract (English)
We search for efficient disentanglers on random Clifford circuits of two-qubit gates arranged in a brick-wall pattern, using the proximal policy optimization (PPO) algorithm \cite{schulman2017proximalpolicyoptimizationalgorithms}. Disentanglers are defined as a set of projective measurements inserted between consecutive entangling layers. An efficient disentangler is a set of projective measurements that minimize the averaged von Neumann entropy of the final state with the least number of total projections possible. The problem is naturally amenable to reinforcement learning techniques by taking the binary matrix representing the projective measurements along the circuit as our state, and actions as bit flipping operations on this binary matrix that add or delete measurements at specified locations. We give rewards to our agent dependent on the averaged von Neumann entropy of the final state and the configuration of measurements, such that the agent learns the optimal policy that will take him from the initial state of no measurements to the optimal measurement state that minimizes the entanglement entropy. Our results indicate that the number of measurements required to disentangle a random quantum circuit is drastically less than the numerical results of measurement-induced phase transition papers. Additionally, the reinforcement learning procedure enables us to characterize the pattern of optimal disentanglers, which is not possible in the works of measurement-induced phase transitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。