无需信道状态信息,用定位数据实现毫米波波束聚焦的多智能体强化学习框架
Learning to Reflect: Hierarchical Multi-Agent Reinforcement Learning for CSI-Free mmWave Beam-Focusing
- 分层多智能体强化学习,高阶分配用户,低阶优化波束焦点
- 相比集中式方案提升2.81-7.94 dB接收信号强度,复杂度越高优势越明显
- 不依赖信道估计,对定位误差和反射面尺寸变化均具鲁棒性
可重构智能表面有望重塑无线环境,但实际部署受限于信道状态信息(CSI)估计的高昂开销以及集中式优化带来的维度爆炸。本文提出一种用于毫米波(mmWave)系统中机械可重构反射面控制的分层多智能体强化学习(HMARL)框架。引入“无CSI”范式,以现成的用户定位数据替代基于导频的信道估计。为应对巨大的组合动作空间,采用基于中心化训练、去中心化执行(CTDE)范式的多智能体近端策略优化(MAPPO)。该架构将控制问题分解为两个抽象层次:高层控制器负责用户与反射面的分配,低层控制器则进行局部波束聚焦优化。全面的射线追踪评估表明,该框架在性能上相较集中式基线提升2.81–7.94 dB RSSI,且系统复杂度越高优势越显著。可扩展性分析显示,当用户密度翻倍时,系统仍保持高效,每用户性能下降极小,总功耗稳定。此外,鲁棒性验证表明其在反射面孔径(45–99个单元)变化下表现良好,并在定位误差达0.5米时仍能实现平滑性能退化。通过消除CSI开销而维持高保真波束聚焦,本工作确立了HMARL在智能毫米波环境中的实用价值。
原文摘要 · Abstract (English)
Reconfigurable Intelligent Surfaces promise to transform wireless environments, yet practical deployment is hindered by the prohibitive overhead of Channel State Information (CSI) estimation and the dimensionality explosion inherent in centralized optimization. This paper proposes a Hierarchical Multi-Agent Reinforcement Learning (HMARL) framework for the control of mechanically reconfigurable reflective surfaces in millimeter-wave (mmWave) systems. We introduce a "CSI-free" paradigm that substitutes pilot-based channel estimation with readily available user localization data. To manage the massive combinatorial action space, the proposed architecture utilizes Multi-Agent Proximal Policy Optimization (MAPPO) under a Centralized Training with Decentralized Execution (CTDE) paradigm. The proposed architecture decomposes the control problem into two abstraction levels: a high-level controller for user-to-reflector allocation and decentralized low-level controllers for low-level focal point optimization. Comprehensive ray-tracing evaluations demonstrate that the framework achieves 2.81-7.94 dB RSSI improvements over centralized baselines, with the performance advantage widening as system complexity increases. Scalability analysis reveals that the system maintains sustained efficiency, exhibiting minimal per-user performance degradation and stable total power utilization even when user density doubles. Furthermore, robustness validation confirms the framework's viability across varying reflector aperture sizes (45-99 tiles) and demonstrates graceful performance degradation under localization errors up to 0.5 m. By eliminating CSI overhead while maintaining high-fidelity beam-focusing, this work establishes HMARL as a practical solution for intelligent mmWave environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。