arXiv:2409.15595cs.AIeess.SP2024-09被引 1

融合物理模型与强化学习,提升自动驾驶车队在延迟下的安全跟驰能力。

Physics Enhanced Residual Policy Learning (PERPL) for safety cruising in mixed traffic platooning under actuator and communication delay

  • 用物理模型保证稳定可解释性,残差策略学习动态调整适应环境变化。
  • 在真实车流和极端条件下,头距误差更小,振荡抑制效果优于传统方法。
  • 适合关注自动驾驶车队控制、鲁棒性设计的工程师与研究者。

线性控制模型因简单易用且支持稳定性分析而广泛应用于车辆控制,但缺乏对动态环境和多目标场景的适应能力。强化学习虽具备自适应性,却存在可解释性差和泛化能力不足的问题。本文提出物理增强残差策略学习(PERPL)框架,结合物理模型的数据高效性与可解释性,以及强化学习对多目标的灵活性和快速计算优势。该框架中,物理部分提供稳定性和可解释性,基于学习的残差策略则对物理策略进行微调以适应环境变化,从而优化决策。我们将其应用于采用恒定时间间距(CTG)策略的混合交通车队(含联网自动驾驶车辆与人工驾驶车辆)的去中心化控制,并考虑执行器与通信延迟。实验表明,在人为极端条件及真实前车轨迹下,所提方法相比线性模型和纯强化学习,显著减小了头距误差并有效抑制振荡。宏观层面,随着采用PERPL方案的自动驾驶车辆渗透率提升,整体交通振荡也明显降低。

原文摘要 · Abstract (English)

Linear control models have gained extensive application in vehicle control due to their simplicity, ease of use, and support for stability analysis. However, these models lack adaptability to the changing environment and multi-objective settings. Reinforcement learning (RL) models, on the other hand, offer adaptability but suffer from a lack of interpretability and generalization capabilities. This paper aims to develop a family of RL-based controllers enhanced by physics-informed policies, leveraging the advantages of both physics-based models (data-efficient and interpretable) and RL methods (flexible to multiple objectives and fast computing). We propose the Physics-Enhanced Residual Policy Learning (PERPL) framework, where the physics component provides model interpretability and stability. The learning-based Residual Policy adjusts the physics-based policy to adapt to the changing environment, thereby refining the decisions of the physics model. We apply our proposed model to decentralized control to mixed traffic platoon of Connected and Automated Vehicles (CAVs) and Human-driven Vehicles (HVs) using a constant time gap (CTG) strategy for cruising and incorporating actuator and communication delays. Experimental results demonstrate that our method achieves smaller headway errors and better oscillation dampening than linear models and RL alone in scenarios with artificially extreme conditions and real preceding vehicle trajectories. At the macroscopic level, overall traffic oscillations are also reduced as the penetration rate of CAVs employing the PERPL scheme increases.

自动驾驶强化学习车队控制物理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。