arXiv:2411.08360cs.LGeess.SP2024-11被引 2

提出新方法提升多环境强化学习的数据覆盖,显著降低误差并加速训练。

Coverage Analysis for Digital Cousin Selection -- Improving Multi-Environment Q-Learning

  • 基于覆盖率系数分析,建立上下界以评估不同环境价值。
  • 新算法使平均策略误差降低65%,速度比穷举搜索快95%。
  • 适合需要高效稳定强化学习的复杂网络优化场景。

Q-learning广泛应用于未知动态的大维度网络优化。近期提出的多环境混合Q-learning(MEMQ)算法在多个结构相关但不同的环境中并行运行多个独立Q-learning,相比多种先进算法在准确性、复杂度和鲁棒性上表现更优。本文开展全面的概率覆盖分析,推导出MEMQ算法中不同覆盖率系数(CC)的期望与方差上下界,据此提出一种简单高效的环境优劣比较方法,性能接近此前提出的部分排序法。进一步设计了一种基于覆盖率的新MEMQ算法,提升了现有MEMQ的准确性和效率。通过四类不同图特性的随机网络图进行数值实验,该算法相较部分排序法降低65%的平均策略误差(APE),且比穷举搜索快95%;相比若干先进强化学习及前期MEMQ算法,其APE降低60%。同时验证了理论结果的正确性,并展示了随动作空间增大仍具可扩展性。

原文摘要 · Abstract (English)

Q-learning is widely employed for optimizing various large-dimensional networks with unknown system dynamics. Recent advancements include multi-environment mixed Q-learning (MEMQ) algorithms, which utilize multiple independent Q-learning algorithms across multiple, structurally related but distinct environments and outperform several state-of-the-art Q-learning algorithms in terms of accuracy, complexity, and robustness. We herein conduct a comprehensive probabilistic coverage analysis to ensure optimal data coverage conditions for MEMQ algorithms. First, we derive upper and lower bounds on the expectation and variance of different coverage coefficients (CC) for MEMQ algorithms. Leveraging these bounds, we develop a simple way of comparing the utilities of multiple environments in MEMQ algorithms. This approach appears to be near optimal versus our previously proposed partial ordering approach. We also present a novel CC-based MEMQ algorithm to improve the accuracy and complexity of existing MEMQ algorithms. Numerical experiments are conducted using random network graphs with four different graph properties. Our algorithm can reduce the average policy error (APE) by 65% compared to partial ordering and is 95% faster than the exhaustive search. It also achieves 60% less APE than several state-of-the-art reinforcement learning and prior MEMQ algorithms. Additionally, we numerically verify the theoretical results and show their scalability with the action-space size.

强化学习Q-learning多环境网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。