arXiv:2608.02826quant-phcs.AI2026-08

量子算法加速强化学习策略求解,逼近理论最优性能。

Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model

  • 融合量子均值估计算法与经典优化技术,设计新型量子迭代方法。
  • 在有限时域和无限时域任务中,查询复杂度优于已有量子方案。
  • 适合对量子机器学习、强化学习效率有追求的研究者参考。

强化学习研究智能体如何与环境交互以最大化回报。标准框架为马尔可夫决策过程(MDPs),目标是寻找最优策略——指导智能体选择动作的函数。本文研究两类MDP:有限时域与无限时域折扣情形,提出新的量子算法以计算近似最优策略。算法结合标准值迭代与量子子程序(如量子均值估计、量子最大值查找),并引入样本最优的经典算法技巧。整体查询复杂度优于先前工作,接近已知的量子下界。

原文摘要 · Abstract (English)

Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible. A standard approach to study such interaction is through Markov Decision Processes (MDPs) and the task of choosing an optimal policy --- a function that tells the agent which action to take. In this work, we study two types of MDPs --- finite-horizon and infinite-horizon discounted --- and propose new quantum algorithms for computing approximate optimal policies. Our quantum algorithms are based on a new combination of standard value iteration and quantum subroutines like quantum mean estimation and quantum maximum finding, overall enhanced with techniques from sample-optimal classical algorithms. Our resulting query complexities improve upon previous works, thus approaching already established quantum lower bounds.

量子算法强化学习值迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。