arXiv:2604.13098cs.MAcs.CV2026-04中稿 · CVPR

用大模型生成交通协调奖励,让红绿灯更懂人情

C$^2$T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

  • 从大模型提取常识知识,生成智能交通奖励函数
  • 在CityFlow上提升效率、安全性和能耗表现
  • 可灵活调整提示词,实现高效或安全优先策略

当前先进的城市交通控制越来越多地采用多智能体强化学习(MARL)来协调交通信号灯控制器(TLCs)与联网自动驾驶汽车(CAVs)。然而,这些系统的性能受限于人工设计的短期奖励机制(如路口压力),难以捕捉安全、通行稳定性和舒适性等高层次人类目标。为此,我们提出C2T框架,从交通-车辆动态中学习一种基于常识的协调模型。该框架将大型语言模型(LLM)中的“常识”知识提炼为可学习的内在奖励函数,并用于指导基于CityFlow的多交叉口基准上的协作式多交叉口TLC MARL系统。实验表明,该框架在交通效率、安全性及能耗代理指标上显著优于多个强基线。此外,我们还展示了其原理上的灵活性:通过修改LLM提示词,可分别生成侧重效率或侧重安全的策略。

原文摘要 · Abstract (English)

State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (TLCs) and Connected Autonomous Vehicles (CAVs). However, the performance of these systems is fundamentally capped by their hand-crafted, myopic rewards (e.g., intersection pressure), which fail to capture high-level, human-centric goals like safety, flow stability, and comfort. To overcome this limitation, we introduce C2T, a novel framework that learns a common-sense coordination model from traffic-vehicle dynamics. C2T distills "common-sense" knowledge from a Large Language Model (LLM) into a learned intrinsic reward function. This new reward is then used to guide the coordination policy of a cooperative multi-intersection TLC MARL system on CityFlow-based multi-intersection benchmarks. Our framework significantly outperforms strong MARL baselines in traffic efficiency, safety, and an energy-related proxy. We further highlight C2T's flexibility in principle, allowing distinct "efficiency-focused" versus "safety-focused" policies by modifying the LLM prompt.

交通控制多智能体大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。