分布式量子强化学习框架,提升多智能体系统训练效率
MADQRL: Distributed Quantum Reinforcement Learning Framework for Multi-Agent Environments

- 多智能体独立学习,分散联合训练负载
- 在合作乒乓环境中比其他分布式策略提升约10%
- 适合高维多智能体场景,兼顾量子与经典模型优势
强化学习是解决现实应用问题的实用方法,但高维环境使传统算法计算成本高昂。近年来量子计算在编码压缩、表征增强、采样优化等方面取得进展,推动了量子强化学习(QRL)发展。然而当前量子硬件尚难支持复杂多智能体系统。为此,本文提出一种分布式量子强化学习框架,多个智能体在不同设备上独立学习,减轻联合训练压力。该方法适用于动作与观测空间不重叠的环境,也可通过合理近似扩展至其他系统。在合作乒乓环境上的实验表明,相比其他分布式策略性能提升约10%,较经典策略模型提升约5%。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is one of the most practical ways to learn from real-life use-cases. Motivated from the cognitive methods used by humans makes it a widely acceptable strategy in the field of artificial intelligence. Most of the environments used for RL are often high-dimensional, and traditional RL algorithms becomes computationally expensive and challenging to effectively learn from such systems. Recent advancements in practical demonstration of quantum computing (QC) theories, such as compact encoding, enhanced representation and learning algorithms, random sampling, or the inherent stochastic nature of quantum systems, have opened up new directions to tackle these challenges. Quantum reinforcement learning (QRL) is seeking significant traction over the past few years. However, the current state of quantum hardware is not enough to cater for such high-dimensional environments with complex multi-agent setup. To tackle this issue, we propose a distributed framework for QRL where multiple agents learn independently, distributing the load of joint training from individual machines. Our method works well for environments with disjoint sets of action and observation spaces, but can also be extended to other systems with reasonable approximations. We analyze the proposed method on cooperative-pong environment and our results indicate ~10% improvement from other distribution strategies, and ~5% improvement from classical models of policy representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。