提出并行混合动作空间强化学习模型,同步优化信号相位与持续时间。
A Parallel Hybrid Action Space Reinforcement Learning Model for Real-world Adaptive Traffic Signal Control
- 设计并行混合动作空间,同时输出相位选择与持续时间参数。
- 在真实交通数据集上减少平均通行时间18.3%,提升通行效率。
- 适合智能交通系统研究者与城市交通管理者参考。
自适应交通信号控制(ATSC)可通过动态调整信号时序有效降低车辆通行时间,但在实时动态不确定交通条件下面临决策复杂性的挑战。本文提出一种并行混合动作空间强化学习模型(PH-DDPG),可同步优化交通信号的相位与持续时间,避免传统两阶段模型的顺序决策。该模型采用面向交通控制的任务特定并行混合动作空间,直接联合输出离散相位选择及其连续持续时间参数,通过统一参数优化实现对动态交通的内在适应。为进一步验证方法的有效性与鲁棒性,我们进行了消融实验,重点考察了在评论家网络中使用随机动作参数掩码的效果,该策略解耦各动作的参数空间,便于为每个动作选用最优参数。实验结果表明,该方法显著提升了真实场景下的适用性与性能表现。
原文摘要 · Abstract (English)
Adaptive traffic signal control (ATSC) can effectively reduce vehicle travel times by dynamically adjusting signal timings but poses a critical challenge in real-world scenarios due to the complexity of real-time decision-making in dynamic and uncertain traffic conditions. The burgeoning field of intelligent transportation systems, bolstered by artificial intelligence techniques and extensive data availability, offers new prospects for the implementation of ATSC. In this study, we introduce a parallel hybrid action space reinforcement learning model (PH-DDPG) that optimizes traffic signal phase and duration of traffic signals simultaneously, eliminating the need for sequential decision-making seen in traditional two-stage models. Our model features a task-specific parallel hybrid action space tailored for adaptive traffic control, which directly outputs discrete phase selections and their associated continuous duration parameters concurrently, thereby inherently addressing dynamic traffic adaptation through unified parametric optimization. %Our model features a unique parallel hybrid action space that allows for the simultaneous output of each action and its optimal parameters, streamlining the decision-making process. Furthermore, to ascertain the robustness and effectiveness of this approach, we executed ablation studies focusing on the utilization of a random action parameter mask within the critic network, which decouples the parameter space for individual actions, facilitating the use of preferable parameters for each action. The results from these studies confirm the efficacy of this method, distinctly enhancing real-world applicability
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。