用强化学习与多模型辩论优化交通灯控制,提升效率与可解释性。
CuraLight: Debate-Guided Data Curation for LLM-Centered Traffic Signal Control

- RL代理生成高质量交互数据,用于LLM控制器的模仿微调。
- 多模型辩论系统提供偏好监督信号,优化信号配时决策。
- 在三座城市真实路网中,平均通行时间减少5.34%以上。
交通信号控制(TSC)是智能交通系统的核心,旨在降低拥堵、排放和出行时间。基于强化学习(RL)和大语言模型(LLM)的方法虽提升了自适应性,但仍面临可解释性差、交互数据不足及对异构路口泛化能力弱的问题。本文提出CuraLight,一种以LLM为中心的框架:由RL代理探索交通环境并生成高质量交互轨迹,转换为提示-响应对用于模仿微调;再通过多LLM集成辩论系统,对候选信号配时方案进行结构化辩论,提供偏好感知的监督信号。在SUMO仿真中,基于济南、杭州、亦庄三个真实路网的实验表明,CuraLight持续优于现有先进基线,平均通行时间降低5.34%,平均队列长度减少5.14%,平均等待时间下降7.02%。结果验证了结合RL辅助探索与辩论式数据梳理在可扩展、可解释交通控制中的有效性。
原文摘要 · Abstract (English)
Traffic signal control (TSC) is a core component of intelligent transportation systems (ITS), aiming to reduce congestion, emissions, and travel time. Recent approaches based on reinforcement learning (RL) and large language models (LLMs) have improved adaptivity, but still suffer from limited interpretability, insufficient interaction data, and weak generalization to heterogeneous intersections. This paper proposes CuraLight, an LLM-centered framework where an RL agent assists the fine-tuning of an LLM-based traffic signal controller. The RL agent explores traffic environments and generates high-quality interaction trajectories, which are converted into prompt-response pairs for imitation fine-tuning. A multi-LLM ensemble deliberation system further evaluates candidate signal timing actions through structured debate, providing preference-aware supervision signals for training. Experiments conducted in SUMO across heterogeneous real-world networks from Jinan, Hangzhou, and Yizhuang demonstrate that CuraLight consistently outperforms state-of-the-art baselines, reducing average travel time by 5.34 percent, average queue length by 5.14 percent, and average waiting time by 7.02 percent. The results highlight the effectiveness of combining RL-assisted exploration with deliberation-based data curation for scalable and interpretable traffic signal control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。