arXiv:2603.08931cs.NIcs.LG2026-03被引 1

用数字孪生优化强化学习训练,降低通信延迟。

Optimizing Reinforcement Learning Training over Digital Twin Enabled Multi-fidelity Networks

  • 分层强化学习结合对抗损失与PPO优化天线角度
  • 物理网络数据采集延迟降低28.01%以上
  • 适合无线网络优化与数字孪生系统研究者

本文提出一种由数字网络孪生(DNT)辅助的深度学习模型训练框架。在物理网络中,基站(BS)通过多天线服务多个移动用户,需动态调整天线俯仰角以优化用户数据速率。由于用户移动性,基站难以准确追踪信道和移动状态,因此采用强化学习(RL)动态调整俯仰角。训练可使用物理网络和DNT的数据:物理网络数据更准确但通信开销高。为此,本文将问题建模为联合优化俯仰角策略与数据采集策略的优化问题,目标是在控制物理网络数据采集延迟的前提下最大化用户数据速率。提出一种融合鲁棒对抗损失与近端策略优化(PPO)的分层强化学习框架。仿真表明,所提方法相比基线模型,物理网络数据采集延迟降低达28.01%和1倍。

原文摘要 · Abstract (English)

In this paper, we investigate a novel digital network twin (DNT) assisted deep learning (DL) model training framework. In particular, we consider a physical network where a base station (BS) uses several antennas to serve multiple mobile users, and a DNT that is a virtual representation of the physical network. The BS must adjust its antenna tilt angles to optimize the data rates of all users. Due to user mobility, the BS may not be able to accurately track network dynamics such as wireless channels and user mobilities. Hence, a reinforcement learning (RL) approach is used to dynamically adjust the antenna tilt angles. To train the RL, we can use data collected from the physical network and the DNT. The data collected from the physical network is more accurate but incurs more communication overhead compared to the data collected from the DNT. Therefore, it is necessary to determine the ratio of data collected from the physical network and the DNT to improve the training of the RL model. We formulate this problem as an optimization problem whose goal is to jointly optimize the tilt angle adjustment policy and the data collection strategy, aiming to maximize the data rates of all users while constraining the time delay introduced by collecting data from the physical network. To solve this problem, we propose a hierarchical RL framework that integrates robust adversarial loss and proximal policy optimization (PPO). Simulation results show that our proposed method reduces the physical network data collection delay by up to 28.01% and 1x compared to a hierarchical RL that uses vanilla PPO as the first level RL, and the baseline that uses robust-RL at the first level and selects the data collection ratio randomly.

强化学习数字孪生无线网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。