量子自然策略梯度算法显著降低量子强化学习的查询次数。
Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
- 用确定性梯度估计替代随机采样,适配量子系统。
- 量子查询复杂度降至 ε⁻¹·⁵,优于经典下界 ε⁻²。
- 适合研究量子强化学习与高效量子算法设计者。
针对具有量子预言机访问能力的无模型量子强化学习(QRL)问题,本文提出量子自然策略梯度(QNPG)算法。该算法将经典自然策略梯度(NPG)中使用的随机采样替换为确定性梯度估计方法,实现与量子系统的无缝集成。尽管此修改引入了有界偏差,但偏差随截断级别增加呈指数衰减。本文证明,所提QNPG算法在量子预言机查询上的样本复杂度为 $ ilde{ ext{O}}(ε^{-1.5})$,显著优于经典方法对马尔可夫决策过程(MDP)查询的 $ ilde{ ext{O}}(ε^{-2})$ 下界。
原文摘要 · Abstract (English)
We address the problem of quantum reinforcement learning (QRL) under model-free settings with quantum oracle access to the Markov Decision Process (MDP). This paper introduces a Quantum Natural Policy Gradient (QNPG) algorithm, which replaces the random sampling used in classical Natural Policy Gradient (NPG) estimators with a deterministic gradient estimation approach, enabling seamless integration into quantum systems. While this modification introduces a bounded bias in the estimator, the bias decays exponentially with increasing truncation levels. This paper demonstrates that the proposed QNPG algorithm achieves a sample complexity of $\tilde{\mathcal{O}}(ε^{-1.5})$ for queries to the quantum oracle, significantly improving the classical lower bound of $\tilde{\mathcal{O}}(ε^{-2})$ for queries to the MDP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。