用强化学习解决网约车调度难题,兼顾效率与公平。
Scalable Ride-Sourcing Vehicle Rebalancing with Service Accessibility Guarantee: A Constrained Mean-Field Reinforcement Learning Approach
- 基于平均场强化学习,实现大规模车队的高效动态调度。
- 可处理数万辆车,训练时间接近线性规划,且不需重训。
- 兼顾服务覆盖率与响应效率,适合城市交通管理与平台运营者。
网约车平台如Uber和Lyft的普及重塑了城市出行方式,但车辆供需在时空上的失衡带来调度难题。本文提出连续状态下的平均场控制(MFC)与平均场强化学习(MFRL)模型,通过建模车辆与整体分布的交互而非个体间交互,有效缓解维度灾难问题,使大规模车队协调的计算复杂度显著降低,并支持无需重训练的弹性扩展。为保障服务公平性,引入可达性约束,设计出在需求满足率与区域覆盖公平性之间权衡的调度策略。基于深圳数据的仿真验证表明,该方法可扩展至数万辆车辆,训练时间与线性规划相当;在车队利用率、订单完成率、接驾距离等关键指标上优于传统基准,同时确保地理区域间的公平服务覆盖。
原文摘要 · Abstract (English)
The expansion of ride-sourcing services such as Uber and Lyft has reshaped urban transportation by offering flexible, on-demand mobility via mobile applications. Despite convenience, these platforms confront significant operational challenges, particularly vehicle rebalancing-strategic repositioning of a fleet of vehicles to address spatiotemporal mismatches in supply and demand. Inadequate rebalancing results in prolonged rider waiting times and inefficient vehicle utilization, but also leads to fairness issues, such as the inequitable distribution of service and disparities in driver income. To tackle these, we introduce continuous-state mean-field control (MFC) and mean-field reinforcement learning (MFRL) models with continuous repositioning actions. MFC and MFRL offer scalable solutions by modeling each vehicle's behavior through interaction with the vehicle distribution, rather than with individual vehicles. This mitigates the curse of dimensionality with respect to the number of agents, enabling coordination across large fleets with significantly reduced computational complexity and eliminating the need to retrain the model when fleet size changes. To ensure equitable service access across geographic regions, we integrate an accessibility constraint into models and derive rebalancing policies that strike a balance between high fulfillment of rider demand and fair coverage of vehicle supply. Extensive evaluation using data-driven simulation of Shenzhen demonstrates the efficiency and robustness of our approach. Remarkably, it scales to tens of thousands of vehicles, with training times comparable to linear programming rebalancing. Besides, our policies effectively explore the efficiency-equity Pareto front, outperforming conventional benchmarks across key metrics like fleet utilization, fulfilled requests, and pickup distance, while ensuring equitable service access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。