arXiv:2608.30317cs.LGcs.AI2026-08

用强化学习在线估计交通需求,提升实时精度与效率

Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance

论文配图:Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance
图 1 · 摘自论文原文
  • 引入链路流量传播引导机制,让策略更新更精准
  • 在墨尔本数据上实现RMSE 4.69、相关性0.995的高精度
  • 适合需要快速响应的智能交通系统应用

在线动态起讫矩阵估计(DODE)旨在根据观测到的链路流量轨迹,校准随时间变化的起讫需求。在在线场景中,需基于当前观测与网络状态估计需求,而后续观测和随机动态网络加载(DNL)结果仍不确定。近年来,强化学习(RL)因其可降低计算负担且适用于随机环境而成为有前景的替代方案。然而,传统方法在离线训练后在线部署时,面对不同目标链路流量轨迹,同一需求向量可能需不同调整,导致标量反馈模糊。为此,本文提出LFPG-RL,将链路流量传播引导(LFPG)整合至近端策略优化(PPO)。LFPG结合链路流量误差敏感度与各起讫-时间需求分量对模拟链路流的贡献,将整体不匹配转化为针对起讫需求的优势塑造,用于PPO策略更新。部署时仅需单次前向传播。该方法在基于链路传输模型与随机路径选择的墨尔本主干道网络15分钟链路流量数据上进行测试,覆盖250个工作日轨迹。在留出轨迹上,其RMSE为4.69,MAPE为20.15%,皮尔逊相关系数达0.995,验证了该方法在在线需求校准中的高效性与准确性。

原文摘要 · Abstract (English)

Online dynamic origin-destination (OD) matrix estimation (DODE) calibrates time-dependent OD demand to reproduce observed link-flow trajectories. In online, OD demand should be estimated from current observations and propagated network states while subsequent observations and stochastic dynamic network loading (DNL) outcomes remain uncertain. Recently, reinforcement learning (RL) has emerged as a promising alternative, reducing computational burden by replacing iterative algorithms while being applicable to stochastic environments. However, because the policy is trained offline and deployed online, it must handle varying target link-flow trajectories; since each target trajectory defines the link-flow error used in the reward, the same OD demand vector can require different adjustments, making conventional scalar feedback ambiguous. To address this gap, this study proposes LFPG-RL, which integrates link-flow propagation guidance (LFPG) into proximal policy optimization (PPO). LFPG combines link-flow error sensitivities with the contribution of each OD-time demand component to simulated link flows, transforming aggregate mismatch into OD-specific advantage shaping for PPO actor updates. At deployment, the policy requires only a single forward pass. LFPG-RL is developed and evaluated on 250 weekday trajectories of 15-min link-flow data from a Melbourne arterial network modeled by a link transmission model with stochastic route choice. On held-out trajectories, LFPG-RL achieved an RMSE of 4.69, MAPE of 20.15%, and Pearson correlation of 0.995. These results support the contention that our method is a more efficient and accurate online OD demand calibration method compared to existing ones.

交通预测强化学习在线估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。