通过学习优化策略,解决临时工招聘延迟与替换成本问题。
Sequential Hiring of Contingent Workers Through Learning-Based Optimization

- 基于实时生产数据,分周期动态决定何时换人及换谁。
- 在时间跨度内,算法后悔项逼近理论下界,性能最优。
- 适合有用工延迟和替换成本的灵活用工场景使用。
本文研究了在工人产出和劳动力供给均不确定的临时用工环境下,企业如何通过序列化决策最大化累计利润。企业需维持固定规模团队,同时随时间学习工人生产效率。该问题存在两大运营摩擦:替换工人成本高,且因先前工作安排、排班约束或入职流程,工人无法立即聘用,导致招聘决策存在随机延迟。我们将其建模为带高成本切换和延迟动作的随机多臂赌博机问题,提出一种基于学习的招聘策略DR-UCB(DelayedReplacement-UCB),通过学习周期逐次做出替换与雇佣决策。每轮周期利用实时生产数据判断何时启动人力调整及选择哪些工人替换或招聘。我们证明该策略的主导阶后悔项与其理论下界在时间跨度上的依赖关系一致。数值实验表明,DR-UCB优于基准策略。
原文摘要 · Abstract (English)
In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and labor supply. A firm seeks to maximize cumulative profit by maintaining an active team of fixed size while learning worker productivity over time. We emphasize two critical operational frictions in this problem: replacing workers is costly, and workers may not be available immediately for hiring because of, for example, prior job commitments, scheduling constraints, or onboarding procedures. Thus, hiring decisions take effect only after a random delay. We formulate this problem as a stochastic multi-play bandit with costly switching and delayed actions, and develop a learning-based hiring policy, DR-UCB (DelayedReplacement-UCB), that makes replacement and hiring decisions sequentially through learning cycles. In each cycle, the policy uses real-time production data to determine when to initiate workforce changes and which workers to replace and hire. We show that the leading-order regret of the proposed policy matches its lower bound in its dependence on the time horizon. Our numerical experiments show that DR-UCB outperforms benchmark policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。