arXiv:2506.20543cs.LGmath.OC2025-06

用真实数据验证强化学习路由算法在技能队列中的高效性

Demonstration of effective UCB-based routing in skill-based queues on real-world data

  • 基于真实数据集,用强化学习动态优化客户路由策略
  • 算法自适应环境变化,延迟更低且优于固定规则策略
  • 可兼顾收益、负载均衡与等待时间,适合复杂服务系统

本文研究了数据中⼼、云计算网络和服务系统等技能型队列系统的最优控制问题。通过一个基于真实数据集的案例研究,我们检验了一种近期提出的强化学习算法在客户路由中的实际应用效果。实验表明,该算法能有效学习并适应环境变化,性能优于静态基准策略,具备实际部署潜力。我们还引入一种新启发式路由规则以降低延迟,并证明该算法可同时优化多个目标:除收益最大化外,还可兼顾服务器负载公平性和客户等待时间减少。通过调节参数实现不同目标间的权衡。最后,我们分析了估计误差和参数调优对算法的影响,为复杂真实队列系统中自适应路由算法的实施提供了重要参考。

原文摘要 · Abstract (English)

This paper is about optimally controlling skill-based queueing systems such as data centers, cloud computing networks, and service systems. By means of a case study using a real-world data set, we investigate the practical implementation of a recently developed reinforcement learning algorithm for optimal customer routing. Our experiments show that the algorithm efficiently learns and adapts to changing environments and outperforms static benchmark policies, indicating its potential for live implementation. We also augment the real-world applicability of this algorithm by introducing a new heuristic routing rule to reduce delays. Moreover, we show that the algorithm can optimize for multiple objectives: next to payoff maximization, secondary objectives such as server load fairness and customer waiting time reduction can be incorporated. Tuning parameters are used for balancing inherent performance trade--offs. Lastly, we investigate the sensitivity to estimation errors and parameter tuning, providing valuable insights for implementing adaptive routing algorithms in complex real-world queueing systems.

强化学习队列调度真实数据多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。