arXiv:2504.08074cs.LGcs.SY2025-04被引 2

用强化学习动态调整可交易信用方案中的收费,提升交通效率。

Deep Reinforcement Learning for Day-to-day Dynamic Tolling in Tradable Credit Schemes

  • 将动态收费问题建模为马尔可夫决策过程,用强化学习求解。
  • 算法在不同交通流量下表现稳定,旅行时间和社会福利接近最优基准。
  • 方法具备跨场景迁移能力,适合实际城市交通系统部署。

可交易信用方案(TCS)作为一种新型拥堵收费机制,具有收入中性与公平性优势,但其设计与实施面临用户行为、市场反应及供需动态等挑战。本文聚焦控制机制,针对TCS下的日度动态收费问题,将其建模为离散时间马尔可夫决策过程,并采用强化学习(RL)算法求解。结果表明,所提方法在旅行时间与社会福利方面与贝叶斯优化基准相当,且在不同路网容量和需求水平下具备良好泛化能力。进一步通过正则化技术缓解动作震荡,提升策略稳定性,获得可在日度供需波动下转移使用的实用收费策略。最后讨论了大规模网络扩展的挑战,并提出利用迁移学习提升计算效率,促进基于RL的TCS方案落地应用。

原文摘要 · Abstract (English)

Tradable credit schemes (TCS) are an increasingly studied alternative to congestion pricing, given their revenue neutrality and ability to address issues of equity through the initial credit allocation. Modeling TCS to aid future design and implementation is associated with challenges involving user and market behaviors, demand-supply dynamics, and control mechanisms. In this paper, we focus on the latter and address the day-to-day dynamic tolling problem under TCS, which is formulated as a discrete-time Markov Decision Process and solved using reinforcement learning (RL) algorithms. Our results indicate that RL algorithms achieve travel times and social welfare comparable to the Bayesian optimization benchmark, with generalization across varying capacities and demand levels. We further assess the robustness of RL under different hyperparameters and apply regularization techniques to mitigate action oscillation, which generates practical tolling strategies that are transferable under day-to-day demand and supply variability. Finally, we discuss potential challenges such as scaling to large networks, and show how transfer learning can be leveraged to improve computational efficiency and facilitate the practical deployment of RL-based TCS solutions.

交通优化强化学习政策设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。