arXiv:2502.14741cs.NIcs.LG2025-02被引 1

用图注意力网络提升光路复用下的路由波长分配效率

Reinforcement Learning with Graph Attention for Routing and Wavelength Assignment with Lightpath Reuse

  • 用图注意力网络建模网络拓扑,优化路由与波长分配决策
  • 相比之前最优强化学习方法,吞吐量提升2.5%(17.4 Tbps)
  • 适合研究光网络资源调度或强化学习应用的工程师与学者

现有研究多关注灵活网格网络中的路由与频谱分配,但实际生产系统普遍采用固定网格搭配灵活速率波长转换器。本文针对这一场景下的光路复用路由与波长分配(RWA-LR)问题进行重新审视,系统性地评估启发式算法,发现按跳数而非总长度排序候选路径可使吞吐率提高6%。采用图注意力网络构建策略与价值函数的强化学习代理,在真实网络结构上实现更优决策。我们公开全部代码,并在实验中超越此前最优强化学习方法2.5%(平均额外吞吐17.4 Tbps),优于最佳启发式算法1.2%(平均额外吞吐8.5 Tbps)。该微小增益反映出长时序资源分配任务中学习高效策略的挑战性。

原文摘要 · Abstract (English)

Many works have investigated reinforcement learning (RL) for routing and spectrum assignment on flex-grid networks but only one work to date has examined RL for fixed-grid with flex-rate transponders, despite production systems using this paradigm. Flex-rate transponders allow existing lightpaths to accommodate new services, a task we term routing and wavelength assignment with lightpath reuse (RWA-LR). We re-examine this problem and present a thorough benchmarking of heuristic algorithms for RWA-LR, which are shown to have 6% increased throughput when candidate paths are ordered by number of hops, rather than total length. We train an RL agent for RWA-LR with graph attention networks for the policy and value functions to exploit the graph-structured data. We provide details of our methodology and open source all of our code for reproduction. We outperform the previous state-of-the-art RL approach by 2.5% (17.4 Tbps mean additional throughput) and the best heuristic by 1.2% (8.5 Tbps mean additional throughput). This marginal gain highlights the difficulty in learning effective RL policies on long horizon resource allocation tasks.

强化学习光网络图神经网络资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。