arXiv:2410.07611cs.LGcs.SY2024-10被引 8

用大视觉模型构建数字孪生,让强化学习更高效地应对动态用户网络关联与负载均衡。

Large Vision Model-Enhanced Digital Twin with Deep Reinforcement Learning for User Association and Load Balancing in Dynamic Wireless Networks

  • 基于街景图生成用户移动轨迹,零样本训练数字孪生环境。
  • 在数字孪生中训练强化学习模型,实现接近真实环境的性能,边缘用户性能提升近20%。
  • 适合研究动态无线网络优化、数字孪生与强化学习融合的科研人员。

密集部署的蜂窝网络中用户关联优化因用户移动性和用户数量波动而极具挑战。尽管深度强化学习(DRL)展现出前景,但其实际应用受限于现实世界高试错成本及训练期间物理网络性能不佳。现有基于DRL的方法通常仅适用于固定用户数场景,受收敛与兼容性限制。为此,本文提出一种大视觉模型(LVM)增强的无线网络数字孪生(DT),并设计并行DT驱动的DRL方法,用于处理用户数量、分布和移动模式动态变化的场景。为构建该LVM增强的DT,我们开发了名为Map2Traj的零样本生成式用户移动模型,基于扩散模型仅从街景图估计用户轨迹模式与空间分布。DRL模型在数字孪生环境中训练,避免与物理网络直接交互。为进一步提升DRL模型在动态场景下的泛化能力,引入并行数字孪生框架,缓解单一环境训练中的强相关性和非平稳性问题,提升训练效率。数值结果表明,所提出的LVM增强数字孪生在训练效能上接近真实环境,且并行框架相比单个真实环境在边缘用户性能上提升近20%。

原文摘要 · Abstract (English)

Optimization of user association in a densely deployed cellular network is usually challenging and even more complicated due to the dynamic nature of user mobility and fluctuation in user counts. While deep reinforcement learning (DRL) emerges as a promising solution, its application in practice is hindered by high trial-and-error costs in real world and unsatisfactory physical network performance during training. Also, existing DRL-based user association methods are typically applicable to scenarios with a fixed number of users due to convergence and compatibility challenges. To address these limitations, we introduce a large vision model (LVM)-enhanced digital twin (DT) for wireless networks and propose a parallel DT-driven DRL method for user association and load balancing in networks with dynamic user counts, distribution, and mobility patterns. To construct this LVM-enhanced DT for DRL training, we develop a zero-shot generative user mobility model, named Map2Traj, based on the diffusion model. Map2Traj estimates user trajectory patterns and spatial distributions solely from street maps. DRL models undergo training in the DT environment, avoiding direct interactions with physical networks. To enhance the generalization ability of DRL models for dynamic scenarios, a parallel DT framework is further established to alleviate strong correlation and non-stationarity in single-environment training and improve training efficiency. Numerical results show that the developed LVM-enhanced DT achieves closely comparable training efficacy to the real environment, and the proposed parallel DT framework even outperforms the single real-world environment in DRL training with nearly 20\% gain in terms of cell-edge user performance.

数字孪生强化学习无线网络用户关联

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。