提出可证明性能边界的数字孪生评估方法,解决训练与实际部署不一致问题。
Provable Performance Bounds for Digital Twin-driven Deep Reinforcement Learning in Wireless Networks: A Novel Digital-Twin Bisimulation Metric
- 基于Wasserstein距离设计数字孪生双模拟度量(DT-BSM)
- 证明真实网络性能损失由孪生模型误差和内部策略次优性共同决定
- 提出可采样计算的实证方法,适用于大规模无线网络场景
数字孪生驱动的深度强化学习已成为无线网络优化的有前景范式,提供安全高效的策略探索环境。然而,现有方法无法在部署前保证数字孪生训练策略的真实世界性能,因缺乏通用指标来评估数字孪生支持可靠强化学习迁移的能力。本文提出基于Wasserstein距离的数字孪生双模拟度量(DT-BSM),用于量化数字孪生环境与真实无线网络中马尔可夫决策过程(MDP)之间的差异。我们证明:任意数字孪生训练策略在真实部署中的次优性(悔差)被其在数字孪生中的悔差与DT-BSM的加权和所限制。为降低大尺度无线网络中Wasserstein距离的计算复杂度,进一步引入基于总变差距离的改进型DT-BSM。针对真实环境中难以获取精确转移概率的问题,提出基于统计采样的经验型DT-BSM方法,并证明其收敛性,同时建立样本量与近似精度之间的定量关系。数值实验验证了该理论成果,首次实现了数字孪生驱动强化学习的可证明且可计算的性能边界。
原文摘要 · Abstract (English)
Digital twin (DT)-driven deep reinforcement learning (DRL) has emerged as a promising paradigm for wireless network optimization, offering safe and efficient training environment for policy exploration. However, in theory existing methods cannot always guarantee real-world performance of DT-trained policies before actual deployment, due to the absence of a universal metric for assessing DT's ability to support reliable DRL training transferrable to physical networks. In this paper, we propose the DT bisimulation metric (DT-BSM), a novel metric based on the Wasserstein distance, to quantify the discrepancy between Markov decision processes (MDPs) in both the DT and the corresponding real-world wireless network environment. We prove that for any DT-trained policy, the sub-optimality of its performance (regret) in the real-world deployment is bounded by a weighted sum of the DT-BSM and its sub-optimality within the MDP in the DT. Then, a modified DT-BSM based on the total variation distance is also introduced to avoid the prohibitive calculation complexity of Wasserstein distance for large-scale wireless network scenarios. Further, to tackle the challenge of obtaining accurate transition probabilities of the MDP in real world for the DT-BSM calculation, we propose an empirical DT-BSM method based on statistical sampling. We prove that the empirical DT-BSM always converges to the desired theoretical one, and quantitatively establish the relationship between the required sample size and the target level of approximation accuracy. Numerical experiments validate this first theoretical finding on the provable and calculable performance bounds for DT-driven DRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。