arXiv:2605.03434cs.LGquant-ph2026-05被引 1

用量子电路提升强化学习效率,减少66%参数量

Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits

论文配图:Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits
图 1 · 摘自论文原文
  • 用变分量子电路替代经典模块,构建混合层级强化学习模型
  • 量子特征提取器使性能超越经典基线,参数量减少66%
  • 发现量子选项价值估计是性能瓶颈,提供高效设计指南

强化学习是极具挑战性的学习范式,其效率与效果提升极为重要。层级强化学习通过时间抽象来结构化决策过程。尽管参数化量子计算在非层级强化学习中已取得成功,但这些优势能否适用于层级决策仍是一个关键开放问题。本文提出一种基于选项-批评架构的混合层级智能体,将特征提取器、选项价值函数、终止函数和选项内策略等经典组件替换为变分量子电路。在标准基准环境上的评估表明,使用量子特征提取器的混合智能体可超越经典基线,同时节省高达66%的可训练参数。进一步消融实验揭示了量子电路架构选择对性能的影响,并识别出量子选项价值估计严重降低性能的架构瓶颈。本工作确立了参数高效的混合层级智能体的设计原则。

原文摘要 · Abstract (English)

Reinforcement learning is one of the most challenging learning paradigms where efficacy and efficiency gains are extremely valuable. Hierarchical reinforcement learning is a variant that leverages temporal abstraction to structure decision-making. While parametrized quantum computations have shown success in non-hierarchical reinforcement learning, whether these advantages adapt to hierarchical decision-making remains a critical open question. In this work, we develop a hybrid hierarchical agent based on the option-critic architecture. This hybrid agent substitutes classical components with variational quantum circuits for feature extractors, option-value functions, termination functions, and intra-option policies. Evaluated on standard benchmarking environments, results show that a hybrid agent utilizing a quantum feature extractor can outperform classical baselines while saving up to 66\% trainable parameters. We also identify an architectural bottleneck that quantum option-value estimation severely degrades performance. Further ablation studies reveal how architectural choices of the quantum circuits affect performance. Our work establishes design principles for parameter-efficient hybrid hierarchical agents.

量子强化学习混合模型参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。