用扩散模型动态调度司机补贴,提升网约车供需匹配效率。
D$^3$-Subsidy: Online and Sequential Driver Subsidy Decision-Making for Large-Scale Ride-Hailing Market

- 基于扩散模型生成未来供需轨迹,实现城市级补贴的前瞻决策。
- 在严格补贴上限下,使订单完成率和平台收入显著提升。
- 适合大规模网约车平台实时调度,兼顾响应速度与预算控制。
滴滴等网约车平台面临高度动态的供需平衡挑战。司机端补贴是调节供需、提升订单完成量(Rides)和平台商品价值(GMV)的关键手段,但实际优化需同时满足三大约束:对随机波动的快速响应、严格的补贴率上限、以及城市尺度下的低延迟执行。这排除了逐单优化的高成本方案,亟需一种面向未来的、考虑约束的城市级在线序列决策控制器。为此,本文提出 D³-Subsidy(动态司机端扩散补贴),一个可部署的城市级补贴控制框架。该框架采用前缀条件扩散模型,从历史数据中生成符合现实的未来供需轨迹,确保训练与线上固定历史部署的一致性;再通过上下文条件反向模块将计划解码为低维城市级控制信号。为实现高效执行,引入基于拉格朗日对偶的映射机制,将补贴率上限直接嵌入订单-司机激励中,避免迭代优化。此外,采用多城市预训练与参数高效微调策略,提升跨城市迁移能力。大量离线评估显示,D³-Subsidy 在提升 Rides 与 GMV 的同时保障补贴上限合规;真实世界 A/B 测试验证其显著提升业绩,且预算违规指标维持在运营阈值内。
原文摘要 · Abstract (English)
Ride-hailing platforms like DiDi Chuxing operate in highly dynamic environments where balancing driver supply and passenger demand is critical. Although driver-side subsidies serve as a primary lever to align these forces and improve key KPIs like completed rides (\texttt{Rides}) and gross merchandise value (\texttt{GMV}), optimizing them in production requires simultaneously meeting three constraints: (i) responsiveness to stochastic shocks, (ii) strict subsidy-rate caps, and (iii) low-latency execution at city scale. These requirements rule out expensive per-order optimization, calling for a forward-looking, constraint-aware city-level controller for online sequential decision making. To meet these requirements, we introduce D$^3$-Subsidy (Dynamic Driver-side Diffusion-based Subsidy), a hierarchical diffusion-based framework for deployable city-wide subsidy control. To bridge the train-inference gap, D$^3$-Subsidy employs a prefix-conditioned diffusion model that samples plausible future trajectories from immutable historical observations, ensuring the training protocol aligns with the fixed-history nature of online deployment. These generated plans are then decoded by a context-conditioned inverse module into low-dimensional city-level control signals. For scalable execution, we bridge the gap between city-level planning and fine-grained dispatch via a Lagrangian-dual-derived mapping, which embeds subsidy-rate caps directly into order-driver incentives without iterative optimization. Additionally, a multi-city pretraining strategy with parameter-efficient fine-tuning enables robust transfer across heterogeneous cities. Extensive offline evaluations demonstrate that D$^3$-Subsidy improves \texttt{Rides} and \texttt{GMV} while enhancing cap compliance, and a real-world A/B test confirms significant uplift while keeping budget-related violation metrics within operational thresholds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。