arXiv:2511.17112quant-phcs.LG2025-11被引 3

拆解量子强化学习关键组件,发现量子优势并非必然。

Dissecting Quantum Reinforcement Learning: A Systematic Evaluation of Key Components

  • 用统一框架对比量子与经典组件,控制变量验证效果
  • 输出复用技术虽纯经典但能提升训练稳定性
  • 强纠缠反而降低优化效率,量子优势受限于设计

基于参数化量子电路(PQC)的量子强化学习(QRL)在量子计算与强化学习交叉领域展现出潜力。由于训练不稳、平坦区(BPs)及组件贡献难分离,其实用性仍存疑。本文通过系统实验评估三大关键环节:(i) 数据嵌入策略,以数据重上传(DR)为代表;(ii) 电路结构设计,重点分析纠缠作用;(iii) 量子测量后的后处理模块,聚焦未被充分研究的输出复用(OR)技术。采用统一的PPO-CartPole框架,在相同条件下对比混合与纯经典智能体。结果表明,尽管输出复用为纯经典方法,但在混合架构中表现出独特行为;数据重上传可提升训练可实现性与稳定性;而更强纠缠反而损害优化性能,抵消了经典收益。这些发现提供了量子与经典贡献相互作用的可控实证,建立了可复现的系统性基准与组件级分析框架。

原文摘要 · Abstract (English)

Parameterised quantum circuit (PQC) based Quantum Reinforcement Learning (QRL) has emerged as a promising paradigm at the intersection of quantum computing and reinforcement learning (RL). By design, PQCs create hybrid quantum-classical models, but their practical applicability remains uncertain due to training instabilities, barren plateaus (BPs), and the difficulty of isolating the contribution of individual pipeline components. In this work, we dissect PQC based QRL architectures through a systematic experimental evaluation of three aspects recurrently identified as critical: (i) data embedding strategies, with Data Reuploading (DR) as an advanced approach; (ii) ansatz design, particularly the role of entanglement; and (iii) post-processing blocks after quantum measurement, with a focus on the underexplored Output Reuse (OR) technique. Using a unified PPO-CartPole framework, we perform controlled comparisons between hybrid and classical agents under identical conditions. Our results show that OR, though purely classical, exhibits distinct behaviour in hybrid pipelines, that DR improves trainability and stability, and that stronger entanglement can degrade optimisation, offsetting classical gains. Together, these findings provide controlled empirical evidence of the interplay between quantum and classical contributions, and establish a reproducible framework for systematic benchmarking and component-wise analysis in QRL.

量子强化学习参数化量子电路纠缠效应训练稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。