arXiv:2509.00713quant-phcs.AI2025-09被引 3

用多芯片量子网络玩超级马里奥,突破硬件限制实现稳定强化学习。

It's-A-Me, Quantum Mario: Scalable Quantum Reinforcement Learning with Multi-Chip Ensembles

  • 拆分复杂观测到多个小量子电路,再用经典网络融合结果。
  • 在超级马里奥环境上表现优于经典模型和单芯片量子模型。
  • 适合想在现有量子硬件上尝试强化学习的研究者。

量子强化学习(QRL)有望以紧凑的函数近似器访问巨大的希尔伯特空间,但受制于当前含噪声中等规模量子(NISQ)设备的局限性,如量子比特数量有限和噪声累积。本文提出一种基于多芯片集成框架的方法,利用多个小型量子卷积神经网络(QCNNs)克服这些限制。该方法将超级马里奥游戏环境中高维观测值分割至独立的量子电路处理,并在双深度Q网络(DDQN)框架下对各量子输出进行经典聚合。这种模块化架构使原本无法由量子智能体处理的复杂环境成为可能,相较经典基线和单芯片量子模型,展现出更优性能与学习稳定性。多芯片集成通过减少维度压缩带来的信息损失,实现了可扩展性提升,且可在近中期量子硬件上实现,为将QRL应用于实际问题提供了可行路径。

原文摘要 · Abstract (English)

Quantum reinforcement learning (QRL) promises compact function approximators with access to vast Hilbert spaces, but its practical progress is slowed by NISQ-era constraints such as limited qubits and noise accumulation. We introduce a multi-chip ensemble framework using multiple small Quantum Convolutional Neural Networks (QCNNs) to overcome these constraints. Our approach partitions complex, high-dimensional observations from the Super Mario Bros environment across independent quantum circuits, then classically aggregates their outputs within a Double Deep Q-Network (DDQN) framework. This modular architecture enables QRL in complex environments previously inaccessible to quantum agents, achieving superior performance and learning stability compared to classical baselines and single-chip quantum models. The multi-chip ensemble demonstrates enhanced scalability by reducing information loss from dimensionality reduction while remaining implementable on near-term quantum hardware, providing a practical pathway for applying QRL to real-world problems.

量子强化学习多芯片集成超级马里奥近中期量子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。