用经典数据训练量子环境,实现量子强化学习优势
From Classical Data to Quantum Advantage -- Quantum Policy Evaluation on Quantum Hardware
- 从经典数据中学习量子环境参数,实现端到端量子建模
- 在量子硬件上完成策略评估,展现量子加速潜力
- 适合研究量子机器学习与强化学习交叉的学者
量子策略评估(QPE)是一种强化学习算法,相比经典的蒙特卡洛估计具有二次加速优势。它通过单位算符直接实现有限马尔可夫决策过程的量子化,其中智能体与环境以叠加态交换状态、动作和奖励。此前,量子环境仅在量子模拟器中手工构建用于演示。本文首次展示如何利用量子机器学习(QML)在量子硬件上从一批经典观测数据中学习环境参数。所学量子环境被用于实际的QPE计算,实现量子硬件上的策略评估。实验表明,尽管存在噪声和相干时间短等挑战,该方法在强化学习中展现出实现量子优势的前景。
原文摘要 · Abstract (English)
Quantum policy evaluation (QPE) is a reinforcement learning (RL) algorithm which is quadratically more efficient than an analogous classical Monte Carlo estimation. It makes use of a direct quantum mechanical realization of a finite Markov decision process, in which the agent and the environment are modeled by unitary operators and exchange states, actions, and rewards in superposition. Previously, the quantum environment has been implemented and parametrized manually for an illustrative benchmark using a quantum simulator. In this paper, we demonstrate how these environment parameters can be learned from a batch of classical observational data through quantum machine learning (QML) on quantum hardware. The learned quantum environment is then applied in QPE to also compute policy evaluations on quantum hardware. Our experiments reveal that, despite challenges such as noise and short coherence times, the integration of QML and QPE shows promising potential for achieving quantum advantage in RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。