FitLight让交通信号灯智能体零成本适配新环境,快速上手且低资源运行。
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control
- 通过联邦模仿学习实现无预训练的实时策略迁移
- 仅需数个训练回合即找到最优控制策略,收敛速度提升显著
- 支持微控制器部署,内存占用低至16KB RAM
尽管基于强化学习(RL)的交通信号控制(TSC)方法已被广泛研究,但其实际应用仍面临高学习成本和泛化能力差的问题。这是因为‘试错’式训练使RL智能体高度依赖特定交通环境,且收敛时间长。为此,我们提出一种新型联邦模仿学习(FIL)框架FitLight,用于多交叉口TSC,使RL智能体可即插即用,无需额外预训练成本。与依赖预训练演示的传统模仿学习不同,FitLight支持实时模仿学习并无缝过渡到强化学习。得益于所提出的知识共享机制和新型基于压力的混合智能体设计,智能体仅需少数几轮即可快速找到最优控制策略。此外,在资源受限场景下,FitLight支持模型剪枝与异构模型聚合,使智能体可在仅含16KB RAM和32KB ROM的微控制器上运行。大量实验表明,相比现有最优方法,FitLight不仅提供更优初始起点,且在真实世界与合成数据集上均收敛至更优解,即使在极端资源限制下依然表现良好。
原文摘要 · Abstract (English)
Although Reinforcement Learning (RL)-based Traffic Signal Control (TSC) methods have been extensively studied, their practical applications still raise some serious issues such as high learning cost and poor generalizability. This is because the ``trial-and-error'' training style makes RL agents extremely dependent on the specific traffic environment, which also requires a long convergence time. To address these issues, we propose a novel Federated Imitation Learning (FIL)-based framework for multi-intersection TSC, named FitLight, which allows RL agents to plug-and-play for any traffic environment without additional pre-training cost. Unlike existing imitation learning approaches that rely on pre-training RL agents with demonstrations, FitLight allows real-time imitation learning and seamless transition to reinforcement learning. Due to our proposed knowledge-sharing mechanism and novel hybrid pressure-based agent design, RL agents can quickly find a best control policy with only a few episodes. Moreover, for resource-constrained TSC scenarios, FitLight supports model pruning and heterogeneous model aggregation, such that RL agents can work on a micro-controller with merely 16{\it KB} RAM and 32{\it KB} ROM. Extensive experiments demonstrate that, compared to state-of-the-art methods, FitLight not only provides a superior starting point but also converges to a better final solution on both real-world and synthetic datasets, even under extreme resource limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。