无需中心服务器,通过稀疏通信实现6G基站协同资源调度。
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

- 各基站本地运行智能体,仅共享压缩后的全局价值函数参数。
- 通信开销降低76%,在干扰强场景下吞吐量和用户体验最优。
- 适合开放解耦的6G无线接入网,尤其适用于大规模多入多出系统。
本文提出FedCritic-MIMO,一种面向开放解耦6G无线接入网中独立部署小区级控制器的无服务器联邦多智能体强化学习框架。控制器不共享训练器,保留本地执行器与个性化价值网络,仅交换兼容的共享价值函数参数。该框架针对复用因子为1的大规模多输入多输出正交频分多址部署,联合优化用户调度、流级功率分配、波束成形、干扰管理与长期服务质量。每个基站本地执行其执行器,无需集中训练或执行器联邦;价值知识通过干扰感知图进行点对点交换。通过无线感知事件触发、自适应逐层top-k稀疏交换及误差反馈、均衡的干扰感知融合机制实现协作。在固定策略、冻结目标价值回归模型下,建立了平衡压缩的点对点价值递归的有限时间平稳性与一致性保证。在强干扰耦合的复用因子1仿真中,相比启发式、独立学习、集中训练及通信消融基线,FedCritic-MIMO在性能-通信权衡上表现最佳:达到最高外部测试吞吐量,改善用户速率分布与平均信干噪比,提升服务质量满足度,单位传输比特的干扰成本最低。相比未压缩分布式价值交换,其价值函数通信开销减少76%。结果表明,通过兼容共享价值参数的无服务器交换,可在不依赖中心轨迹收集或参数服务器聚合的情况下协调RAN控制器。
原文摘要 · Abstract (English)
This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term QoS with limited inter-controller signaling. Each base station locally executes its actor without centralized training or actor federation, while critic knowledge is exchanged peer-to-peer over an interference-aware graph. It enables this collaboration through wireless-aware event triggering, adaptive layer-wise top-$k$ sparse critic exchange with error feedback, and balanced interference-aware fusion. We establish conditional finite-time stationarity and consensus guarantees for the balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. It achieves the highest held-out throughput, improves user-rate distribution and mean SINR, increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It reduces critic-communication overhead by $76\%$ relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate RAN controllers without centralized trajectory collection or parameter-server aggregation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。