arXiv:2607.09422cs.LGcond-mat.mes-hall2026-07

用分解动作空间的方法,让多智能体协作调校量子器件更稳定高效。

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

论文配图:Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning
图 1 · 摘自论文原文
  • 通过在线学习动作空间分解,减少智能体间干扰
  • 零样本泛化到未见规模的量子器件,收敛步数基本恒定
  • 适合需要快速校准大规模量子处理器的研究者

协同多智能体强化学习适用于参数空间大且具有局部可利用结构的问题,如静电定义的量子点阵列调校。然而,若参数交叉干扰强烈,单个智能体视角下的非平稳环境会破坏学习过程——这正是此类系统人工调校时面临的难题。本文提出一种在线学习的动作空间因子化表示,以解耦智能体并最小化其相互干扰。所提出的框架QADAPT利用该因子化策略,基于局部测量和奖励高效学习共享策略。采用此模块化方法,系统实现对未见量子器件尺寸的零样本泛化,并保持接近恒定的收敛步数以达到目标工作区。本研究为大规模量子处理器的快速校准提供了一条可扩展路径。

原文摘要 · Abstract (English)

Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter cross-talk is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning - the same effect that plagues manual tuning of such systems. We propose using a factored representation of the action space, learned online, to decouple agents and minimize their interference. Our framework, QADAPT, uses this factorization to efficiently learn shared policies based on local measurements and rewards. With this modular strategy, we achieve zero-shot generalization to unseen quantum device sizes and maintain an approximately constant number of convergence steps to reach target regimes. This work provides a scalable route toward the rapid calibration of large-scale quantum processors.

量子计算强化学习多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。