arXiv:2603.05000cs.LGcs.MA2026-03

多运营商竞争下,用强化学习优化定价与车辆调度

Competitive Multi-Operator Reinforcement Learning for Joint Pricing and Fleet Rebalancing in AMoD Systems

  • 构建双运营商并行学习框架,结合用户选择理论模拟真实竞争
  • 竞争环境下价格更低,车队布局更分散,行为模式显著区别于垄断
  • 方法对对手策略的不确定性有鲁棒性,适合多主体协同决策场景

自主出行即服务(AMoD)系统有望通过提供高性价比的按需交通服务,满足日益增长的出行需求。然而,现实中的AMoD市场将是多个运营商竞争的格局,他们通过战略定价和车队部署争夺乘客。尽管强化学习在单运营商控制优化中表现良好,但现有研究未能捕捉市场竞争动态。本文提出一种多运营商强化学习框架,让两个运营商同时学习定价与车队再平衡策略。通过引入离散选择理论,使乘客分配与需求竞争可内生地从效用最大化决策中产生。利用多个城市的实际数据进行实验表明,竞争从根本上改变了学习到的行为模式:相比垄断情形,价格更低,车队定位模式也明显不同。值得注意的是,基于学习的方法对竞争带来的额外随机性具有鲁棒性,即使无法完全观测对手策略,竞争主体仍能成功收敛到有效策略。

原文摘要 · Abstract (English)

Autonomous Mobility-on-Demand (AMoD) systems promise to revolutionize urban transportation by providing affordable on-demand services to meet growing travel demand. However, realistic AMoD markets will be competitive, with multiple operators competing for passengers through strategic pricing and fleet deployment. While reinforcement learning has shown promise in optimizing single-operator AMoD control, existing work fails to capture competitive market dynamics. We investigate the impact of competition on policy learning by introducing a multi-operator reinforcement learning framework where two operators simultaneously learn pricing and fleet rebalancing policies. By integrating discrete choice theory, we enable passenger allocation and demand competition to emerge endogenously from utility-maximizing decisions. Experiments using real-world data from multiple cities demonstrate that competition fundamentally alters learned behaviors, leading to lower prices and distinct fleet positioning patterns compared to monopolistic settings. Notably, we demonstrate that learning-based approaches are robust to the additional stochasticity of competition, with competitive agents successfully converging to effective policies while accounting for partially unobserved competitor strategies.

强化学习城市交通多智能体定价策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。