用量子纠缠提升多智能体协作能力,无需通信也能实现更强策略。
Learning to Coordinate via Quantum Entanglement in Multi-Agent Reinforcement Learning
- 通过可微量子测量参数化,让智能体共享量子纠缠来协调决策。
- 在单轮博弈中学习到超越经典随机性的量子优势策略。
- 适用于无通信的多智能体序列决策问题,如Dec-POMDP。
多智能体强化学习(MARL)中缺乏通信严重制约协作能力。以往工作通过共享随机性或相关设备来关联局部策略,以辅助去中心化决策。本文首次提出利用共享量子纠缠作为协调资源的训练框架,其允许的通信无约束相关策略类比于仅使用共享随机性更为广泛。该思路源于量子物理中的经典结论:对于某些无通信的单轮合作博弈,共享量子纠缠能产生优于仅依赖共享随机性的策略,即存在量子优势。我们的框架基于一种新颖的可微策略参数化方法,支持对量子测量进行优化,并采用新策略架构将联合策略分解为量子协调者与去中心化本地执行者。为验证有效性,我们首先证明可在完全黑盒的单轮博弈中,仅凭经验学习出实现量子优势的策略;随后展示了该机制在典型多智能体顺序决策问题——即去中心化部分可观测马尔可夫决策过程(Dec-POMDP)中,仍可学习出具备量子优势的策略。
原文摘要 · Abstract (English)
The inability to communicate poses a major challenge to coordination in multi-agent reinforcement learning (MARL). Prior work has explored correlating local policies via shared randomness, sometimes in the form of a correlation device, as a mechanism to assist in decentralized decision-making. In contrast, this work introduces the first framework for training MARL agents to exploit shared quantum entanglement as a coordination resource, which permits a larger class of communication-free correlated policies than shared randomness alone. This is motivated by well-known results in quantum physics which posit that, for certain single-round cooperative games with no communication, shared quantum entanglement enables strategies that outperform those that only use shared randomness. In such cases, we say that there is quantum advantage. Our framework is based on a novel differentiable policy parameterization that enables optimization over quantum measurements, together with a novel policy architecture that decomposes joint policies into a quantum coordinator and decentralized local actors. To illustrate the effectiveness of our proposed method, we first show that we can learn, purely from experience, strategies that attain quantum advantage in single-round games that are treated as black box oracles. We then demonstrate how our machinery can learn policies with quantum advantage in an illustrative multi-agent sequential decision-making problem formulated as a decentralized partially observable Markov decision process (Dec-POMDP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。