arXiv:2411.06650quant-phcs.LG2024-11

将量子核方法引入强化学习,提升策略梯度效率。

Quantum Policy Gradient in Reproducing Kernel Hilbert Space

  • 用量子核定义策略,实现数据驱动的函数表达
  • 相比经典方法,查询复杂度降低至二次方级
  • 适用于多维动作空间,适合量子环境下的强化学习

参数化量子电路为机器学习提供了高表达力且数据高效的表现形式。由于量子态存在于高维希尔伯特空间,参数化量子电路天然具备核方法的解释。尽管量子核在量子监督学习中已广泛研究,但在量子强化学习中仍被忽视。本文提出使用核策略与量子策略梯度算法,用于可量子访问的环境。在讨论此类策略特性并演示经典策略梯度在量子环境中的相干策略后,我们提出了基于量子核策略的参数化与非参数化策略梯度及演员-评论家算法。该方法结合数值与解析量子策略梯度技术,可利用核方法的优势,如函数及其梯度的数据驱动形式和可调表达力。所提方法适用于向量值动作空间,各形式均比经典对应方法查询复杂度降低为二次方。此外,基于随机策略梯度、确定性策略梯度与自然策略梯度的演员-评论家算法,在有利条件下进一步降低了查询复杂度。

原文摘要 · Abstract (English)

Parametrised quantum circuits offer expressive and data-efficient representations for machine learning. Due to quantum states residing in a high-dimensional Hilbert space, parametrised quantum circuits have a natural interpretation in terms of kernel methods. The representation of quantum circuits in terms of quantum kernels has been studied widely in quantum supervised learning, but has been overlooked in the context of quantum RL. This paper proposes the use of kernel policies and quantum policy gradient algorithms for quantum-accessible environments. After discussing the properties of such policies and a demonstration of classical policy gradient on a coherent policy in a quantum environment, we propose parametric and non-parametric policy gradient and actor-critic algorithms with quantum kernel policies in quantum environments. This approach, implemented with both numerical and analytical quantum policy gradient techniques, allows exploiting the many advantages of kernel methods, including data-driven forms for functions (and their gradients) as well as tunable expressiveness. The proposed approach is suitable for vector-valued action spaces and each of the formulations demonstrates a quadratic reduction in query complexity compared to their classical counterparts. We propose actor-critic algorithms based on stochastic policy gradient, deterministic policy gradient, and natural policy gradient, and demonstrate additional query complexity reductions compared to quantum policy gradient algorithms under favourable conditions.

量子强化学习核方法策略梯度量子计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。