arXiv:2511.06229cs.LG2025-11中稿 · publication in Tra…

用强化学习动态估计交通流量起止点矩阵,解决车辆路径归属难题。

Deep Reinforcement Learning for Dynamic Origin-Destination Matrix Estimation in Microscopic Traffic Simulations Considering Credit Assignment

  • 将交通起止点估计建模为马尔可夫决策过程,通过强化学习逐步生成最优路径矩阵。
  • 在真实高速路网中,相比传统方法降低链路流量均方误差59.2%至88.3%。
  • 适合需要高精度交通仿真校准的智能交通系统研究者使用。

本文聚焦微观交通仿真中的动态起止点矩阵估计(DODE)问题,这是实现有效仿真的关键校准步骤。由于车辆行为具有复杂的时序动态和内在不确定性,难以精确判断每辆车在何时经过哪条路段,导致起止点矩阵与链路流量之间的贡献关系复杂且模糊,形成核心的信用分配难题。为此,本文将DODE问题建模为马尔可夫决策过程(MDP),提出一种基于无模型深度强化学习(DRL)的新框架。该框架中,智能体通过与仿真环境直接交互,学习最优策略以顺序生成起止点矩阵并持续优化。在阮根-杜普瓦网络的简化实验和涵盖圣克拉拉与圣何塞的实路网案例中进行评估。结果表明,所提方法显著优于最强传统基线,在简化实验中降低链路流量均方误差23.7%,在真实案例中降低59.2%-88.3%。通过将DODE重构为序列决策问题,本方法借助学习策略解决信用分配挑战,为微观交通仿真校准提供新范式。

原文摘要 · Abstract (English)

This paper focuses on dynamic origin-destination matrix estimation (DODE), a crucial calibration process necessary for the effective application of microscopic traffic simulations. The fundamental challenge of the DODE problem in microscopic simulations stems from the complex temporal dynamics and inherent uncertainty of individual vehicle dynamics. This makes it highly challenging to precisely determine which vehicle traverses which link at any given moment, resulting in intricate and often ambiguous relationships between origin-destination (OD) matrices and their contributions to resultant link flows. This phenomenon constitutes the credit assignment problem, a central challenge addressed in this study. We formulate the DODE problem as a Markov Decision Process (MDP) and propose a novel framework that applies model-free deep reinforcement learning (DRL). Within our proposed framework, the agent learns an optimal policy to sequentially generate OD matrices, refining its strategy through direct interaction with the simulation environment. This approach was evaluated through a toy experiment on the Nguyen-Dupuis network and a case study utilizing an actual highway subnetwork spanning Santa Clara and San Jose. Experimental results show that the proposed method consistently improves calibration performance relative to the strongest conventional baseline, reducing link-flow MSE by 23.7% in the toy experiment and by 59.2-88.3% in the real-world case study. By reframing DODE as a sequential decision-making problem, our approach addresses the credit assignment challenge through a learned policy and provides a novel framework for calibration of microscopic traffic simulations.

交通仿真强化学习路径估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。