用图神经网络优化无线网络中的分布式采样,降低估计误差。
Decentralized Learning Strategies for Estimation Error Minimization with Graph Neural Networks
- 基于图神经网络的多智能体强化学习框架,实现去中心化策略优化。
- 在多跳无线网络中,性能优于现有方法,且随节点增多效果更佳。
- 策略可迁移至结构相似的大规模网络,对非平稳环境有强鲁棒性。
我们研究动态但结构相似的多跳无线网络中自回归马尔可夫源的实时采样与估计问题。每个节点缓存其他节点的数据并通过无线冲突信道通信,目标是通过去中心化策略最小化时间平均估计误差。由于动作空间维度高、网络拓扑复杂,解析求解最优策略不可行。为此,我们提出一种基于图的多智能体强化学习框架用于策略优化。理论上,我们证明所提策略具备可迁移性,即在某一图上训练的策略可有效应用于结构相似的图。数值实验表明:(i) 所提策略优于当前最优基线;(ii) 训练策略可迁移至更大网络,性能随智能体数量增加而提升;(iii) 图形化训练过程能抵御非平稳性,即使使用独立学习技术亦然;(iv) 循环机制在独立学习和集中训练-分布式执行中均至关重要,显著提升对非平稳性的鲁棒性。
原文摘要 · Abstract (English)
We address real-time sampling and estimation of autoregressive Markovian sources in dynamic yet structurally similar multi-hop wireless networks. Each node caches samples from others and communicates over wireless collision channels, aiming to minimize time-average estimation error via decentralized policies. Due to the high dimensionality of action spaces and complexity of network topologies, deriving optimal policies analytically is intractable. To address this, we propose a graphical multi-agent reinforcement learning framework for policy optimization. Theoretically, we demonstrate that our proposed policies are transferable, allowing a policy trained on one graph to be effectively applied to structurally similar graphs. Numerical experiments demonstrate that (i) our proposed policy outperforms state-of-the-art baselines; (ii) the trained policies are transferable to larger networks, with performance gains increasing with the number of agents; (iii) the graphical training procedure withstands non-stationarity, even when using independent learning techniques; and (iv) recurrence is pivotal in both independent learning and centralized training and decentralized execution, and improves the resilience to non-stationarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。