arXiv:2503.02961cs.LGcs.IT2025-03被引 3

用数学工具分析强化学习泛化能力,解决无线通信中的算法评估难题

Koopman-Based Generalization of Deep Reinforcement Learning With Application to Wireless Communications

  • 基于柯普曼算子建模强化学习的状态动作演化过程
  • 通过谱特征与H∞范数定量评估算法泛化性能
  • 在无人机毫米波通信场景中对比了SAC与PPO的泛化能力

深度强化学习(DRL)是推动无线通信等工程领域发展的关键技术,但其可解释性与泛化能力有限。传统信息论方法因训练数据非独立同分布而不适用于DRL泛化分析。本文提出一种新方法:将训练后DRL算法的状态与动作演化视为未知的离散、随机、非线性动力学函数,并采用数据驱动的柯普曼算子识别方法对其进行逼近,构建两种可解释表示。基于这些表示,建立基于柯普曼算子谱特征分析与H∞范数的严格数学框架,用于评估DRL算法泛化能力。该方法被应用于无人机辅助毫米波无线通信场景,对广泛认可的稳健算法SAC与近端策略优化(PPO)进行泛化性能对比。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) is a key machine learning technology driving progress across various scientific and engineering fields, including wireless communication. However, its limited interpretability and generalizability remain major challenges. In supervised learning, generalizability is commonly evaluated through the generalization error using information-theoretic methods. In DRL, the training data is sequential and not independent and identically distributed (i.i.d.), rendering traditional information-theoretic methods unsuitable for generalizability analysis. To address this challenge, this paper proposes a novel analytical method for evaluating the generalizability of DRL. Specifically, we first model the evolution of states and actions in trained DRL algorithms as unknown discrete, stochastic, and nonlinear dynamical functions. Then, we employ a data-driven identification method, the Koopman operator, to approximate these functions, and propose two interpretable representations. Based on these interpretable representations, we develop a rigorous mathematical approach to evaluate the generalizability of DRL algorithms. This approach is formulated using the spectral feature analysis of the Koopman operator, leveraging the H_\infty norm. Finally, we apply this generalization analysis to compare the soft actor-critic method, widely recognized as a robust DRL approach, against the proximal policy optimization algorithm for an unmanned aerial vehicle-assisted mmWave wireless communication scenario.

强化学习泛化分析柯普曼算子无线通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。