用模拟增强强化学习,让拼车系统不再只看眼前,长远优化匹配与调度。
Non-myopic Matching and Rebalancing in Large-Scale On-Demand Ride-Pooling Systems Using Simulation-Informed Reinforcement Learning
- 在强化学习中嵌入拼车模拟,实现不短视的匹配决策
- 匹配策略提升服务率8.4%,车队规模可减少25%以上
- 加入车辆调度后,乘客等待时间减少27.3%,适合运营优化者
拼车服务可通过共享行程降低乘客成本、减少拥堵和环境影响,但其决策常过于短视,忽视长期影响。本文提出一种基于模拟的强化学习方法,将拼车仿真嵌入学习机制,实现非短视的匹配与调度决策。通过n步时序差分学习处理模拟经验,获得时空状态价值,并利用纽约市出租车请求数据验证效果。结果表明,非短视匹配策略可使服务率提升8.4%,同时减少乘客车内与等待时间;相比短视策略,车队规模可减少超过25%而性能不变,显著降低成本。引入车辆重平衡策略后,等待时间减少27.3%,车内时间减少12.5%,服务率提升15.1%,代价是每乘客行驶里程略增。
原文摘要 · Abstract (English)
Ride-pooling, also known as ride-sharing, shared ride-hailing, or microtransit, is a service wherein passengers share rides. This service can reduce costs for both passengers and operators and reduce congestion and environmental impacts. A key limitation, however, is its myopic decision-making, which overlooks long-term effects of dispatch decisions. To address this, we propose a simulation-informed reinforcement learning (RL) approach. While RL has been widely studied in the context of ride-hailing systems, its application in ride-pooling systems has been less explored. In this study, we extend the learning and planning framework of Xu et al. (2018) from ride-hailing to ride-pooling by embedding a ride-pooling simulation within the learning mechanism to enable non-myopic decision-making. In addition, we propose a complementary policy for rebalancing idle vehicles. By employing n-step temporal difference learning on simulated experiences, we derive spatiotemporal state values and subsequently evaluate the effectiveness of the non-myopic policy using NYC taxi request data. Results demonstrate that the non-myopic policy for matching can increase the service rate by up to 8.4% versus a myopic policy while reducing both in-vehicle and wait times for passengers. Furthermore, the proposed non-myopic policy can decrease fleet size by over 25% compared to a myopic policy, while maintaining the same level of performance, thereby offering significant cost savings for operators. Incorporating rebalancing operations into the proposed framework cuts wait time by up to 27.3%, in-vehicle time by 12.5%, and raises service rate by 15.1% compared to using the framework for matching decisions alone at the cost of increased vehicle minutes traveled per passenger.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。