用强化学习同时优化送餐调度与骑手引导,提升效率与公平性。
Real-Time Integrated Dispatching and Idle Fleet Steering with Deep Reinforcement Learning for A Meal Delivery Platform
- 构建双控强化学习框架,联合优化派单与骑手调度策略。
- 实时决策下配送效率提升,骑手负载更均衡,缺人情况缓解。
- 适合需要高时效、公平性保障的即时配送平台参考。
为实现高服务质量和盈利能力,像Uber Eats和Grubhub这样的外卖平台需战略性运营车队,确保当前订单及时送达,同时避免因决策不当导致未来骑手人力不足。本文提出一种基于强化学习(RL)的战略双控框架,解决外卖平台的实时订单派发与空闲骑手引导问题。针对问题的时序特性,将派单与骑手引导建模为马尔可夫决策过程,并通过深度强化学习(DRL)框架训练策略,输入包含显式预测需求。在双控框架中,派单与引导策略迭代联合训练,生成具有前瞻性的实时决策,在局部与网络层面协同优化。为提升派单公平性,引入卷积深度Q网络构建公平骑手嵌入;为平衡供需,采用均值场近似供给-需求知识,在局部层级重新分配空闲骑手。实验表明,该框架显著提升了配送效率与骑手工作负载公平性,缓解了服务网络中的供不应求现象。本研究为外卖平台及其他按需服务的前瞻性实时运营提供了强化学习解决方案。
原文摘要 · Abstract (English)
To achieve high service quality and profitability, meal delivery platforms like Uber Eats and Grubhub must strategically operate their fleets to ensure timely deliveries for current orders while mitigating the consequential impacts of suboptimal decisions that leads to courier understaffing in the future. This study set out to solve the real-time order dispatching and idle courier steering problems for a meal delivery platform by proposing a reinforcement learning (RL)-based strategic dual-control framework. To address the inherent sequential nature of these problems, we model both order dispatching and courier steering as Markov Decision Processes. Trained via a deep reinforcement learning (DRL) framework, we obtain strategic policies by leveraging the explicitly predicted demands as part of the inputs. In our dual-control framework, the dispatching and steering policies are iteratively trained in an integrated manner. These forward-looking policies can be executed in real-time and provide decisions while jointly considering the impacts on local and network levels. To enhance dispatching fairness, we propose convolutional deep Q networks to construct fair courier embeddings. To simultaneously rebalance the supply and demand within the service network, we propose to utilize mean-field approximated supply-demand knowledge to reallocate idle couriers at the local level. Utilizing the policies generated by the RL-based strategic dual-control framework, we find the delivery efficiency and fairness of workload distribution among couriers have been improved, and under-supplied conditions have been alleviated within the service network. Our study sheds light on designing an RL-based framework to enable forward-looking real-time operations for meal delivery platforms and other on-demand services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。