arXiv:2605.08516cs.AI2026-05

用智能提示和不确定性正则化,让交通信号系统既高效又可解释。

OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control

论文配图:OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control
图 1 · 摘自论文原文
  • 通过设定奖励门槛过滤弱信号,提升强化学习稳定性。
  • 使旅行时间减少75%,队列长度下降67%,且能跨路口直接迁移。
  • 适合关注交通优化与可解释AI的工程师与城市规划者。

透明决策对交通信号控制(TSC)系统赢得公众信任至关重要。传统基于强化学习的TSC方法缺乏可解释性,如同黑箱。尽管大语言模型(LLMs)能提供自然语言推理,但其在TSC中的微调仍不稳定,因反馈稀疏且延迟,多数操作对拥堵指标影响微小。我们提出OracleTSC,通过两项机制稳定基于LLM的TSC:(1) 奖励门槛机制,从环境奖励中减去校准阈值以过滤弱学习信号;(2) 不确定性正则化,最大化所选响应的概率,促进采样输出间的一致性决策。在LibSignal基准上的实验表明,OracleTSC使小型的LLaMA3-8B模型显著提升交通效率:相比预训练基线,旅行时间减少75%,队列长度下降67%,同时通过自然语言解释保持可解释性。该方法还展现出强跨路口泛化能力:在未经过额外微调的情况下,一个路口训练的策略在结构不同的另一路口实现旅行时间降低17%、队列长度降低39%。结果表明,感知不确定性的奖励设计可有效提升强化微调在TSC中的稳定性和效果。

原文摘要 · Abstract (English)

Transparent decision-making is essential for traffic signal control (TSC) systems to earn public trust. However, traditional reinforcement learning-based TSC methods function as black boxes with limited interpretability. Although large language models (LLMs) can provide natural language reasoning, reinforcement finetuning for TSC remains unstable because feedback is sparse and delayed, while most actions produce only marginal changes in congestion metrics. We introduce OracleTSC, which stabilizes LLM-based TSC through two mechanisms: (1) a reward hurdle mechanism that filters weak learning signals by subtracting a calibrated threshold from environmental rewards, and (2) uncertainty regularization that maximizes the probability of the selected response to encourage consistent decisions across sampled outputs. Experiments on the LibSignal benchmark show that OracleTSC enables a compact LLaMA3-8B model to substantially improve traffic efficiency, achieving a 75% reduction in travel time and a 67% decrease in queue length compared with the pretrained baseline while preserving interpretability through natural language explanations. OracleTSC also demonstrates strong cross-intersection generalization: a policy trained on one intersection transfers to a structurally different intersection with 17% lower travel time and 39% lower queue length without additional finetuning. These results suggest that uncertainty-aware reward shaping can improve the stability and effectiveness of reinforcement fine-tuning for TSC.

交通信号可解释AI强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。