用强化学习同时优化人行横道布局与信号灯控制,提升通行效率。
DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning

- 分两阶段:先生成交叉口布局,再学自适应信号控制策略
- 行人到最近路口时间减少23%,车流等待时间降低65%
- 策略可泛化至新需求和不同布局,无需重新训练
现代视觉系统能大规模检测、追踪并预测城市中各类行为体,但感知结果难以转化为实际城市设计。本文提出DeCoR,一种基于强化学习的两阶段框架,利用流量观测数据协同优化人行横道布局与网络级信号控制。设计阶段将行人网络建模为图结构,学习生成式策略,以高斯混合模型参数化人行横道位置与宽度,并从中采样新布局;每个布局下,共享控制策略学习自适应信号配时,以最小化行人与车辆总延迟。在一段750米的真实城市路段上,基于视频与Wi-Fi日志感知需求,DeCoR学习的布局使行人到达最近交叉口的时间减少23%,且使用交叉口数量少于现有配置。控制方面,相比固定周期信号,行人与车辆等待时间分别降低79%和65%。此外,该控制策略对训练外的需求具有泛化能力,且在布局变化下无需重训仍保持鲁棒性。
原文摘要 · Abstract (English)
Modern vision systems can detect, track, and forecast urban actors at scale, yet translating perception outputs to urban design remains limited. We introduce DeCoR, a two-stage reinforcement learning framework that leverages flow observations to co-optimize crosswalk layout and network-level signal control. The design stage encodes the pedestrian network as a graph and learns a generative policy that parameterizes a Gaussian mixture model over crosswalk location and width, from which new crosswalks are sampled. For each layout, a shared control policy learns adaptive signal timings to minimize joint pedestrian and vehicle delay. On a 750 m real-world urban corridor with demand sensed from video and Wi-Fi logs, DeCoR learns a layout that reduces pedestrian arrival time to their nearest crosswalk by 23% while using fewer crosswalks than existing configurations. On the control side, DeCoR reduces pedestrian and vehicle wait time by 79% and 65%, respectively, relative to fixed-time signalization. Further, the control policy generalizes to demands outside of training and is robust to layout changes without retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。