arXiv:2507.11865cs.LG2025-07

优化司机接单策略,提升平台折扣订单匹配效率。

Optimizing Drivers' Discount Order Acceptance Strategies: A Policy-Improved Deep Deterministic Policy Gradient Framework

  • 用改进的深度确定性策略梯度框架动态调整司机接折扣单比例。
  • 实验显示该方法早期训练损失显著降低,学习效率更高。
  • 适合面临数据少、需快速上线的出行平台决策系统。

平台整合迅速发展,通过将多个网约车平台整合至单一应用以缓解市场碎片化问题。为应对乘客偏好差异,第三方集成商推出折扣快送服务,由快递司机以更低价格接单。对单个平台而言,鼓励更多司机参与该服务可扩大需求池并提升匹配效率,但会压缩利润空间。本文从单个平台视角出发,动态管理司机接受折扣订单的策略。由于新商业模式缺乏历史数据,需依赖在线学习,但初期试错成本高昂,亟需可靠的早期性能表现。为此,本文将司机接单比例决策建模为连续控制任务。针对第三方集成商高随机性及匹配机制不透明的问题,提出一种改进的深度确定性策略梯度(pi-DDPG)框架,引入精炼模块提升早期训练阶段的策略性能。基于真实数据构建定制化仿真器验证方法有效性。数值实验表明,pi-DDPG在学习效率上表现更优,显著降低早期训练损失,更适用于实际网约车场景。

原文摘要 · Abstract (English)

The rapid expansion of platform integration has emerged as an effective solution to mitigate market fragmentation by consolidating multiple ride-hailing platforms into a single application. To address heterogeneous passenger preferences, third-party integrators provide Discount Express service delivered by express drivers at lower trip fares. For the individual platform, encouraging broader participation of drivers in Discount Express services has the potential to expand the accessible demand pool and improve matching efficiency, but often at the cost of reduced profit margins. This study aims to dynamically manage drivers' acceptance of Discount Express from the perspective of an individual platform. The lack of historical data under the new business model necessitates online learning. However, early-stage exploration through trial and error can be costly in practice, highlighting the need for reliable early-stage performance in real-world deployment. To address these challenges, this study formulates the decision regarding the proportion of drivers accepting discount orders as a continuous control task. In response to the high stochasticity and the opaque matching mechanisms employed by third-party integrator, we propose an innovative policy-improved deep deterministic policy gradient (pi-DDPG) framework. The proposed framework incorporates a refiner module to boost policy performance during the early training phase. A customized simulator based on a real-world dataset is developed to validate the effectiveness of the proposed pi-DDPG. Numerical experiments demonstrate that pi-DDPG achieves superior learning efficiency and significantly reduces early-stage training losses, enhancing its applicability to practical ride-hailing scenarios.

网约车强化学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。