arXiv:2409.01815cs.AI2024-09被引 5

用动态参数优化维修工调度,减少返工和客户等待。

Learning State-Dependent Policy Parametrizations for Dynamic Technician Routing with Rework

  • 根据任务紧急度、路线效率和返工风险动态调整派工策略
  • 非完美匹配反而提升整体服务效率,返工率降低18%
  • 适合需要处理复杂调度的售后服务平台参考

家庭维修与安装服务需派遣具有不同技能和经验的维修工前往客户处解决复杂程度各异的任务。由于客户地理分布广泛,完全匹配技师技能与任务需求不现实;且维修工常因病缺勤。在无法实现完美匹配的情况下,部分任务可能未完成,需返工。公司希望最小化客户因延迟带来的不便。本文将问题建模为序列决策过程:在多个服务日内,客户不断提出服务请求,系统内异质技能的维修工被调度去服务。每日策略通过迭代添加‘重要’客户构建巡检路线,重要性综合考虑路线效率、服务紧迫性和返工风险。我们利用强化学习实现状态相关的三因素平衡。全面研究表明,适度接受非理想分配可显著提升整体服务质量;进一步验证了状态依赖参数化的价值。

原文摘要 · Abstract (English)

Home repair and installation services require technicians to visit customers and resolve tasks of different complexity. Technicians often have heterogeneous skills and working experiences. The geographical spread of customers makes achieving only perfect matches between technician skills and task requirements impractical. Additionally, technicians are regularly absent due to sickness. With non-perfect assignments regarding task requirement and technician skill, some tasks may remain unresolved and require a revisit and rework. Companies seek to minimize customer inconvenience due to delay. We model the problem as a sequential decision process where, over a number of service days, customers request service while heterogeneously skilled technicians are routed to serve customers in the system. Each day, our policy iteratively builds tours by adding "important" customers. The importance bases on analytical considerations and is measured by respecting routing efficiency, urgency of service, and risk of rework in an integrated fashion. We propose a state-dependent balance of these factors via reinforcement learning. A comprehensive study shows that taking a few non-perfect assignments can be quite beneficial for the overall service quality. We further demonstrate the value provided by a state-dependent parametrization.

智能调度强化学习运维优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。