arXiv:2410.10212cs.AIcs.LG2024-10被引 12

用大模型自动生成奖励函数,让公交调度更智能高效。

Large Language Model-Enhanced Reinforcement Learning for Generic Bus Holding Control Strategies

  • 用大模型自动生成并迭代优化强化学习的奖励函数
  • 在多线路多站点场景下表现优于传统方法20%以上
  • 适合交通系统优化、智能调度等实际应用

公交停站控制是维持公交系统稳定性和提升运营效率的常用策略。传统基于模型的方法常因公交状态预测和乘客需求估计精度低而受限。相比之下,强化学习(RL)作为数据驱动方法,在制定公交停站策略方面展现出巨大潜力。但将现实中稀疏且延迟的控制目标转化为强化学习所需的密集实时奖励极具挑战,通常需大量人工试错。为此,本文提出一种利用大语言模型(LLM)上下文学习与推理能力的自动奖励生成范式。该范式称为大模型增强型强化学习(LLM-enhanced RL),包含奖励初始化器、奖励修正器、性能分析器和奖励精炼器等多个模块,协同完成奖励函数的初始化与迭代优化,并通过过滤无效奖励确保强化学习代理性能稳定演化。为验证其可行性,该方法在不同线路数、站点数和乘客需求的广泛公交停站控制场景中进行测试。结果表明,相比基线强化学习、基于大模型的控制器、物理反馈控制器及优化控制器,该范式在性能、泛化能力和鲁棒性上均具显著优势。

原文摘要 · Abstract (English)

Bus holding control is a widely-adopted strategy for maintaining stability and improving the operational efficiency of bus systems. Traditional model-based methods often face challenges with the low accuracy of bus state prediction and passenger demand estimation. In contrast, Reinforcement Learning (RL), as a data-driven approach, has demonstrated great potential in formulating bus holding strategies. RL determines the optimal control strategies in order to maximize the cumulative reward, which reflects the overall control goals. However, translating sparse and delayed control goals in real-world tasks into dense and real-time rewards for RL is challenging, normally requiring extensive manual trial-and-error. In view of this, this study introduces an automatic reward generation paradigm by leveraging the in-context learning and reasoning capabilities of Large Language Models (LLMs). This new paradigm, termed the LLM-enhanced RL, comprises several LLM-based modules: reward initializer, reward modifier, performance analyzer, and reward refiner. These modules cooperate to initialize and iteratively improve the reward function according to the feedback from training and test results for the specified RL-based task. Ineffective reward functions generated by the LLM are filtered out to ensure the stable evolution of the RL agents' performance over iterations. To evaluate the feasibility of the proposed LLM-enhanced RL paradigm, it is applied to extensive bus holding control scenarios that vary in the number of bus lines, stops, and passenger demand. The results demonstrate the superiority, generalization capability, and robustness of the proposed paradigm compared to vanilla RL strategies, the LLM-based controller, physics-based feedback controllers, and optimization-based controllers. This study sheds light on the great potential of utilizing LLMs in various smart mobility applications.

公交调度强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。