arXiv:2503.01069cs.AIcs.MA2025-03被引 2

用多智能体强化学习优化人力调度,兼顾短期派工与长期人才管理。

Multi-Agent Reinforcement Learning with Long-Term Performance Objectives for Service Workforce Optimization

  • 构建模块化模拟器,统一建模派工、人员管理与定位三者关系
  • 支持可配置的动态场景,涵盖随机性与非平稳性变化
  • 提供基准方法,便于算法对比与消融实验

人力资源优化在组织高效运作中至关重要,决策需跨越多个管理与时间尺度。例如,在即时服务请求调度的同时,还需规划具备不同技能的人才招聘,形成高度动态的优化问题。现有研究多聚焦于资源分配、设施选址等子问题,通常采用局部搜索等启发式方法,近年也引入深度强化学习。然而,这些方法难以准确反映真实场景中各子问题并非完全独立的情况。本文旨在填补这一空白,设计一个统一的人力资源优化模拟器。该模拟器支持多智能体强化学习方法的开发,重点关注人员派工、人力管理与人员定位三个相互关联的方面。通过可配置参数,模拟器能探索具有不同随机性与非平稳性的动态场景。为便于基准测试与消融分析,还集成了启发式与强化学习基线方法。

原文摘要 · Abstract (English)

Workforce optimization plays a crucial role in efficient organizational operations where decision-making may span several different administrative and time scales. For instance, dispatching personnel to immediate service requests while managing talent acquisition with various expertise sets up a highly dynamic optimization problem. Existing work focuses on specific sub-problems such as resource allocation and facility location, which are solved with heuristics like local-search and, more recently, deep reinforcement learning. However, these may not accurately represent real-world scenarios where such sub-problems are not fully independent. Our aim is to fill this gap by creating a simulator that models a unified workforce optimization problem. Specifically, we designed a modular simulator to support the development of reinforcement learning methods for integrated workforce optimization problems. We focus on three interdependent aspects: personnel dispatch, workforce management, and personnel positioning. The simulator provides configurable parameterizations to help explore dynamic scenarios with varying levels of stochasticity and non-stationarity. To facilitate benchmarking and ablation studies, we also include heuristic and RL baselines for the above mentioned aspects.

多智能体强化学习人力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。