arXiv:2505.20889cs.AI2025-05被引 4

用强化学习让导航系统一步步推荐路线,实现全局交通最优。

Reinforcement Learning-based Sequential Route Recommendation for System-Optimal Traffic Assignment

  • 将交通分配问题转为单智能体强化学习任务,逐条推荐路线。
  • 在Braess网络收敛到理论最优解,OW网络偏差仅0.35%。
  • 基于系统最优设计的路线集可加速学习并提升效果,适合交通规划研究者。

现代导航系统与共享出行平台越来越依赖个性化路线推荐以提升个体出行体验和运营效率。然而,一个核心问题仍待解答:这种顺序性、个性化的路线决策能否共同导向系统最优(SO)交通分配?本文提出一种基于学习的框架,将静态的系统最优交通分配问题重构为单智能体深度强化学习任务。一个中心智能体在出行需求到达时,依次为旅行者推荐路线,以最小化整体系统行程时间。为提升学习效率与解的质量,我们设计了融合传统交通分配迭代结构的MSA引导深度Q-learning算法。该方法在Braess和Ortuzar-Willumsen(OW)网络上进行评估,结果表明:在Braess网络中,强化学习智能体成功收敛至理论最优解;在OW网络中,仅产生0.35%的偏差。进一步消融实验显示,路线动作集的设计显著影响收敛速度与最终性能,基于系统最优信息构建的路线集能加快学习并获得更优结果。本工作提供了一种理论严谨且实际可行的学习驱动式序列分配方法,实现了个体路由行为与系统级效率之间的桥梁。

原文摘要 · Abstract (English)

Modern navigation systems and shared mobility platforms increasingly rely on personalized route recommendations to improve individual travel experience and operational efficiency. However, a key question remains: can such sequential, personalized routing decisions collectively lead to system-optimal (SO) traffic assignment? This paper addresses this question by proposing a learning-based framework that reformulates the static SO traffic assignment problem as a single-agent deep reinforcement learning (RL) task. A central agent sequentially recommends routes to travelers as origin-destination (OD) demands arrive, to minimize total system travel time. To enhance learning efficiency and solution quality, we develop an MSA-guided deep Q-learning algorithm that integrates the iterative structure of traditional traffic assignment methods into the RL training process. The proposed approach is evaluated on both the Braess and Ortuzar-Willumsen (OW) networks. Results show that the RL agent converges to the theoretical SO solution in the Braess network and achieves only a 0.35% deviation in the OW network. Further ablation studies demonstrate that the route action set's design significantly impacts convergence speed and final performance, with SO-informed route sets leading to faster learning and better outcomes. This work provides a theoretically grounded and practically relevant approach to bridging individual routing behavior with system-level efficiency through learning-based sequential assignment.

强化学习交通优化路径推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。