简单启发式算法在光网络资源分配中表现优于多数强化学习方案。
Reinforcement Learning for Dynamic Resource Allocation in Optical Networks: Hype or Hope?
- 系统评估多种启发式算法在不同网络拓扑下的表现。
- 发现路径计数与排序规则显著影响性能,简单方法阻塞率低一个数量级。
- 揭示现有强化学习提升空间有限,仅可增载19%-36%。
近年来,强化学习(RL)在光网络动态资源分配中的应用备受关注,相关同行评审论文近100篇。本文综述该领域进展,指出基准测试与可复现性方面存在重大缺口。我们系统评估多种启发式算法在多样网络拓扑下的表现,发现路径选择中的路径计数与排序规则显著影响基准性能。通过严谨重构五篇里程碑论文的问题并应用改进后的基准,结果表明:简单启发式算法在多数情况下表现不亚于甚至超越已发表的强化学习方案,阻塞概率常降低一个数量级。此外,我们提出一种基于去碎片化的新型方法,实证得出网络阻塞的下界,显示在相同阻塞水平下,对基准启发式算法的潜在提升幅度仅为19%-36%的流量负载增加。相关仿真框架与结果已公开,以推动可复现研究与标准化评估(https://doi.org/10.5281/zenodo.12594495)。
原文摘要 · Abstract (English)
The application of reinforcement learning (RL) to dynamic resource allocation in optical networks has been the focus of intense research activity in recent years, with almost 100 peer-reviewed papers. We present a review of progress in the field, and identify significant gaps in benchmarking practices and reproducibility. To determine the strongest benchmark algorithms, we systematically evaluate several heuristics across diverse network topologies. We find that path count and sort criteria for path selection significantly affect the benchmark performance. We meticulously recreate the problems from five landmark papers and apply the improved benchmarks. Our comparisons demonstrate that simple heuristics consistently match or outperform the published RL solutions, often with an order of magnitude lower blocking probability. Furthermore, we present empirical lower bounds on network blocking using a novel defragmentation-based method, revealing that potential improvements over the benchmark heuristics are limited to 19-36% increased traffic load for the same blocking performance in our examples. We make our simulation framework and results publicly available to promote reproducible research and standardized evaluation https://doi.org/10.5281/zenodo.12594495.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。