用多模态大模型让交通信号灯零样本应对突发情况
ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

- 融合传感器、摄像头和预训练控制器,动态判断是否调整信号
- 突发应急车辆时等待时间减少88.7%,且日常通行效率不变
- 适合需要快速响应真实世界突发事件的智能交通系统
强化学习(RL)在交通信号控制(TSC)中表现良好,但依赖预定义状态,难以响应训练数据外的开放世界事件。物联网交叉口提供来自路侧传感器和摄像头的异构观测数据,为提升RL适应性提供了可能。为此,我们提出ReasonLight,一种基于多模态基础模型增强的强化学习框架,实现零样本交通信号控制。ReasonLight整合三类信息:结构化交通测量数据、多视角相机图像与预训练RL控制器的候选相位决策。给定RL提议的相位,ReasonLight从多视角图像中提取视觉语义,并与传感器生成的场景描述进行对齐。该对齐使语义引导的精修模块可依据交通规则与事件语义,决定保留或调整原动作。为确保运行可靠性,精修动作受限于可用相位集合;无效决策将被拒绝,系统回退至原始RL动作。我们在两类训练中未见的罕见事件上评估ReasonLight:应急车辆优先通行与临时交通管制。实验结果表明,ReasonLight实现了无需重训练的零样本适应,在降低应急车辆等待时间最高达88.7%的同时,维持了与基准相当的常规交通性能。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to observable open-world events that are absent from training data. IoT-enabled intersections provide heterogeneous observations from roadside sensors and cameras, creating opportunities to improve RL adaptability to such events. To this end, we propose ReasonLight, a multimodal foundation model-enhanced RL framework for zero-shot TSC. ReasonLight integrates three sources of information: structured traffic measurements, multi-view camera observations, and candidate phase decisions from a pre-trained RL controller. Given an RL-proposed phase, ReasonLight extracts visual semantics from multi-view images and aligns them with compact sensor-derived scene descriptions. This alignment enables a semantic-guided refinement module to either preserve or adjust the proposed action according to traffic rules and event semantics. To ensure operational reliability, refined actions are constrained by the set of available phases. Any invalid decision is rejected, and the system falls back to the original RL action. We evaluate ReasonLight on two types of rare events not seen during RL training: emergency vehicle priority and temporary traffic regulation. Experimental results show that ReasonLight achieves zero-shot adaptation without retraining. It reduces emergency vehicle waiting time by up to 88.7% compared with the RL-only backbone while preserving comparable routine traffic performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。