arXiv:2606.08276quant-phcs.ET2026-06

用量子态直接建模随机环境,提升强化学习的表达与适应能力

QnRL: Quantum-Native Reinforcement Learning

论文配图:QnRL: Quantum-Native Reinforcement Learning
图 1 · 摘自论文原文
  • 基于量子叠加与纠缠,在希尔伯特空间中直接建模环境概率分布
  • 实验显示评估得分最高提升82.9%,参数量平均减少94.3%
  • 适合研究量子机器学习与复杂随机系统决策的学者

量子强化学习(QRL)在随机环境中学习有效决策策略方面具有潜力。现有架构通过估计期望回报间接近环境行为,限制了表达力与自适应性。本文提出量子原生强化学习(QnRL),一种分布式强化学习框架,利用量子系统的分布特性,直接将环境随机变量建模为量子态分布。QnRL通过新提出的量子振幅反冲(QuAK)算法,在希尔伯特空间中比较多个叠加分布的第m阶矩的第n次幂,理论证明可从量子生成模型的矩中提炼出条件动作策略分布,并在该空间内优化。该复杂分布组合还提供了经典及经典采样量子模型无法表达的未知环境相关性维度。跨多种环境的实验表明,QnRL在评估得分上最高提升82.9%,平均参数量减少94.3%,对未见观测的期望回报估计更准确,且对随机条件变化的适应性更强。

原文摘要 · Abstract (English)

Quantum reinforcement learning (QRL) is a promising approach to learn effective decision strategies across several applications with stochastic environments. Instead of directly modeling the random variables that govern these environments, existing QRL architectures indirectly approximate environment behavior by estimating expected outcomes, which limits their expressive power and adaptive potential. Overcoming such challenges requires a novel QRL approach that exploits the distributional nature of quantum computers to directly model environment random variables as quantum state distributions. Hence, in this paper, a novel framework dubbed quantum-native reinforcement learning (QnRL) is proposed. QnRL is a distributional RL framework that learns conditional distributions naturally in Hilbert space via superimposed and entangled quantum states. Thus, QnRL can directly model the behavior of stochastic learning environments via the natural properties of quantum systems. QnRL accomplishes this via a novel, proposed quantum amplitude kickback (QuAK) algorithm that enables comparing the $n$-th power of the $m$-th moment of multiple superimposed distributions. It is theoretically proven that a conditional action policy distribution is distilled from the moments of a quantum generative model entirely within Hilbert space via QuAK, and optimized via QnRL. This complex distribution composition is also shown to provide extra dimensions for expressing environment correlations that are unknown to purely classical and classically-sampled quantum distributional models. Experimental results across diverse environments show that QnRL achieves up to $82.9\%$ higher evaluation scores, with up to $94.3\%$ fewer parameters on average, more accurately estimates the expected return for unseen observations, and better adapts to varying stochastic conditions compared to the baseline.

量子学习强化学习分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。