arXiv:2505.19486eess.SYcs.LG2025-05NeurIPS被引 3

用视觉语言模型提升交通灯控制安全与效率

VLMLight: Safety-Critical Traffic Signal Control via Vision-Language Meta-Control and Dual-Branch Reasoning Architecture

  • 结合视觉感知与语言模型,分场景决策交通信号
  • 紧急车辆等待时间减少65%,常规场景延迟低于1%
  • 适合需要高安全性的智能交通系统部署

交通信号控制(TSC)是城市交通的核心挑战,需在效率与安全间实时权衡。现有方法(从规则启发式到强化学习)在复杂动态、安全关键场景中泛化能力不足。本文提出VLMLight框架,融合视觉-语言元控制与双分支推理结构。核心创新是首个基于图像的交通模拟器,支持交叉口多视角视觉感知,使策略能分析车辆类型、运动状态与空间密度等丰富线索。大型语言模型(LLM)作为安全优先的元控制器,根据情况选择快速强化学习策略或结构化推理分支。在后者中,多个LLM代理协作评估交通相位、优先保障应急车辆并验证规则合规性。实验表明,相较于纯强化学习系统,VLMLight可将应急车辆等待时间降低高达65%,同时在标准条件下保持近实时性能,延迟增加不足1%。该方案为下一代智能交通信号控制提供了可扩展、可解释且注重安全的解决方案。

原文摘要 · Abstract (English)

Traffic signal control (TSC) is a core challenge in urban mobility, where real-time decisions must balance efficiency and safety. Existing methods - ranging from rule-based heuristics to reinforcement learning (RL) - often struggle to generalize to complex, dynamic, and safety-critical scenarios. We introduce VLMLight, a novel TSC framework that integrates vision-language meta-control with dual-branch reasoning. At the core of VLMLight is the first image-based traffic simulator that enables multi-view visual perception at intersections, allowing policies to reason over rich cues such as vehicle type, motion, and spatial density. A large language model (LLM) serves as a safety-prioritized meta-controller, selecting between a fast RL policy for routine traffic and a structured reasoning branch for critical cases. In the latter, multiple LLM agents collaborate to assess traffic phases, prioritize emergency vehicles, and verify rule compliance. Experiments show that VLMLight reduces waiting times for emergency vehicles by up to 65% over RL-only systems, while preserving real-time performance in standard conditions with less than 1% degradation. VLMLight offers a scalable, interpretable, and safety-aware solution for next-generation traffic signal control.

交通控制视觉语言模型安全关键强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。