利用网络动态衰减特性,为多智能体系统设计可扩展的谱表示方法。
Scalable spectral representations for multi-agent reinforcement learning in network MDPs
- 基于网络动态指数衰减,构建每个智能体的局部谱表示
- 在连续状态动作空间中实现算法收敛性保证
- 在基准任务上优于通用函数逼近方法,适合大规模多智能体场景
网络马尔可夫决策过程(Network MDPs)是多智能体控制的常用模型,但由于全局状态-动作空间随智能体数量呈指数增长,高效学习面临挑战。本文利用网络动态的指数衰减特性,首次推导出适用于网络MDPs的可扩展谱局部表示,该表示为每个智能体的局部Q函数诱导出一个网络线性子空间。在此基础上,设计了针对连续状态-动作网络MDPs的可扩展算法框架,并提供了算法收敛性的端到端保证。实验验证了该表示方法在两个基准问题上的有效性,表明其在表示局部Q函数方面优于通用函数逼近方法。
原文摘要 · Abstract (English)
Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global state-action space with the number of agents. In this work, utilizing the exponential decay property of network dynamics, we first derive scalable spectral local representations for network MDPs, which induces a network linear subspace for the local $Q$-function of each agent. Building on these local spectral representations, we design a scalable algorithmic framework for continuous state-action network MDPs, and provide end-to-end guarantees for the convergence of our algorithm. Empirically, we validate the effectiveness of our scalable representation-based approach on two benchmark problems, and demonstrate the advantages of our approach over generic function approximation approaches to representing the local $Q$-functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。