arXiv:2601.18419quant-phcs.AI2026-01中稿 · IEEE ICC 2026

量子多智能体通过通信机制实现自发合作,突破经典困境。

Emergent Cooperation in Quantum Multi-Agent Reinforcement Learning Using Communication

  • 设计四种通信协议,让量子Q-learning智能体互传信号
  • 在三个博弈中合作率超80%,最优方案达95%以上
  • 适合研究量子协作或强化学习的学者参考

经典多智能体强化学习中的自发合作已受广泛关注,尤其在序列社会困境(SSDs)背景下。尽管经典方法已展示出合作潜力,但将此类方法拓展至量子多智能体强化学习的研究仍有限,尤其缺乏基于通信的探索。本文引入四种通信机制:互认令牌交换(MATE)、其扩展版本互信激励互认令牌交换(MEDIATE)、同行奖励机制Gifting,以及强化跨智能体学习(RIAL),并在三个典型序列社会困境——重复囚徒困境、重复猎鹿博弈与重复斗鸡博弈中进行评估。实验结果表明,采用带时序差分度量的MATE(MATE_TD)、AutoMATE、MEDIATE-I和MEDIATE-S的方案,在所有博弈中均实现了高水平合作,最高合作率达95%以上,证明通信是量子多智能体强化学习中促成自发合作的有效机制。

原文摘要 · Abstract (English)

Emergent cooperation in classical Multi-Agent Reinforcement Learning has gained significant attention, particularly in the context of Sequential Social Dilemmas (SSDs). While classical reinforcement learning approaches have demonstrated capability for emergent cooperation, research on extending these methods to Quantum Multi-Agent Reinforcement Learning remains limited, particularly through communication. In this paper, we apply communication approaches to quantum Q-Learning agents: the Mutual Acknowledgment Token Exchange (MATE) protocol, its extension Mutually Endorsed Distributed Incentive Acknowledgment Token Exchange (MEDIATE), the peer rewarding mechanism Gifting, and Reinforced Inter-Agent Learning (RIAL). We evaluate these approaches in three SSDs: the Iterated Prisoner's Dilemma, Iterated Stag Hunt, and Iterated Game of Chicken. Our experimental results show that approaches using MATE with temporal-difference measure (MATE\textsubscript{TD}), AutoMATE, MEDIATE-I, and MEDIATE-S achieved high cooperation levels across all dilemmas, demonstrating that communication is a viable mechanism for fostering emergent cooperation in Quantum Multi-Agent Reinforcement Learning.

量子强化学习多智能体合作机制通信协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。