arXiv:2507.18795cs.AI2025-07被引 1

用强化学习优化排队网络路由,抗干扰且可扩展。

Simulation-Driven Reinforcement Learning in Queuing Network Routing Optimization

  • 结合深度确定性策略梯度与动态规划,提升学习稳定性。
  • 在多种突发状况下仍能快速学出高效路由策略。
  • 适合制造与通信系统中的智能调度场景。

本研究提出一种模拟驱动的强化学习框架,用于优化复杂排队网络中的路由决策,重点关注制造与通信应用。针对传统排队方法在动态、不确定环境中的局限性,我们采用深度确定性策略梯度(DDPG)结合动态规划风格(Dyna-DDPG)的鲁棒强化学习方法。框架包含可灵活配置的仿真环境,能模拟多样化的排队场景、中断和不可预测条件。改进的Dyna-DDPG通过分离建模下一状态转移与奖励,显著提升稳定性和样本效率。大量实验与严格评估表明,该框架能快速学习有效路由策略,在扰动下保持稳健性能,并可有效扩展至更大网络规模。此外,强调了良好的软件工程实践,保障框架的可复现性与可维护性,支持实际部署。

原文摘要 · Abstract (English)

This study focuses on the development of a simulation-driven reinforcement learning (RL) framework for optimizing routing decisions in complex queueing network systems, with a particular emphasis on manufacturing and communication applications. Recognizing the limitations of traditional queueing methods, which often struggle with dynamic, uncertain environments, we propose a robust RL approach leveraging Deep Deterministic Policy Gradient (DDPG) combined with Dyna-style planning (Dyna-DDPG). The framework includes a flexible and configurable simulation environment capable of modeling diverse queueing scenarios, disruptions, and unpredictable conditions. Our enhanced Dyna-DDPG implementation incorporates separate predictive models for next-state transitions and rewards, significantly improving stability and sample efficiency. Comprehensive experiments and rigorous evaluations demonstrate the framework's capability to rapidly learn effective routing policies that maintain robust performance under disruptions and scale effectively to larger network sizes. Additionally, we highlight strong software engineering practices employed to ensure reproducibility and maintainability of the framework, enabling practical deployment in real-world scenarios.

强化学习排队网络路由优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。