用量子计算加速强化学习策略评估,提升效率。
Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning
- 将价值函数编码为量子叠加态,实现多状态动作对并行评估
- 理论证明评估步骤样本复杂度呈多项式降低
- 适合关注量子增强机器学习的科研人员
我们提出 Q-Policy,一种混合量子-经典强化学习框架,通过利用量子计算原语数学上加速策略评估与优化。Q-Policy 将价值函数编码为量子叠加态,借助幅值编码和量子并行性,同时评估多个状态-动作对。我们设计了一种量子增强的策略迭代算法,在标准假设下,评估步骤的样本复杂度实现可证明的多项式降低。为验证方法的技术可行性与理论严谨性,我们在小规模离散控制任务的经典模拟上测试了 Q-Policy。受当前硬件与仿真限制,实验聚焦于概念验证行为,而非大规模实证评估。结果表明,Q-Policy 有望成为未来量子设备上可扩展强化学习的理论基础,解决传统方法难以应对的可扩展性挑战。
原文摘要 · Abstract (English)
We propose Q-Policy, a hybrid quantum-classical reinforcement learning (RL) framework that mathematically accelerates policy evaluation and optimization by exploiting quantum computing primitives. Q-Policy encodes value functions in quantum superposition, enabling simultaneous evaluation of multiple state-action pairs via amplitude encoding and quantum parallelism. We introduce a quantum-enhanced policy iteration algorithm with provable polynomial reductions in sample complexity for the evaluation step, under standard assumptions. To demonstrate the technical feasibility and theoretical soundness of our approach, we validate Q-Policy on classical emulations of small discrete control tasks. Due to current hardware and simulation limitations, our experiments focus on showcasing proof-of-concept behavior rather than large-scale empirical evaluation. Our results support the potential of Q-Policy as a theoretical foundation for scalable RL on future quantum devices, addressing RL scalability challenges beyond classical approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。