30亿参数模型让红绿灯学会类人推理,实时调控交通更高效。
Traffic-R1: Reinforced LLMs Bring Human-Like Reasoning to Traffic Signal Control Systems
- 用大模型自我探索+专家引导强化学习,实现类人决策。
- 部署在手机芯片上实时运行,管理超5.5万司机,队列减少5%以上。
- 支持跨路口协同与可解释控制,减轻人工调度负担。
我们提出Traffic-R1,一个30亿参数的交通信号控制基础模型,通过在仿真交通环境中进行自探索和迭代强化学习,结合专家指导,具备类人推理能力。相比传统强化学习和近期基于大模型的方法,Traffic-R1具有三大优势:零样本泛化能力,无需重训练即可适应新路网和分布外事件;紧凑的30亿参数设计,可在移动级芯片上实现实时推理,支持边缘部署;可解释的控制过程,通过通信与异步网络实现多路口协调。大量基准测试显示,Traffic-R1优于强基线及训练密集的强化学习控制器。实际应用中,该模型已管理每日影响超过5.5万名司机的信号系统,平均排队长度减少超5%,操作员工作量减半。模型开源地址:https://huggingface.co/Season998/Traffic-R1。
原文摘要 · Abstract (English)
We introduce Traffic-R1, a 3B-parameter foundation model with human-like reasoning for Traffic signal control (TSC), developed via self-exploration and iterative reinforcement of LLM with expert guidance in a simulated traffic environment. Compared with traditional reinforcement learning and recent LLM-based methods, Traffic-R1 offers three main advantages: zero-shot generalization, transferring unchanged to new road networks and out-of-distribution incidents by leveraging internal traffic-control policies and reasoning; a compact 3B-parameter design that supports real-time inference on mobile-class chips for edge deployment; and an explainable TSC process that enables multi-intersection coordination through communication and an asynchronous communication network. Extensive benchmarks show Traffic-R1 outperforms strong baselines and training-intensive RL controllers. In production, the model now manages signals affecting over 55,000 drivers daily, reduces average queue lengths by more than 5%, and halves operator workload. Our model is available at https://huggingface.co/Season998/Traffic-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。