arXiv:2409.15105cs.AIcs.MA2024-09被引 5

用Transformer提升自动驾驶车辆协同决策效率与安全性

SPformer: A Transformer Based DRL Decision Making Method for Connected Automated Vehicles

  • 基于Transformer架构,用可学习策略令牌建模多车联合策略
  • 在仿真中实现高效且安全的驾驶决策,优于现有DRL方法
  • 适合研究智能交通系统与多智能体强化学习的学者

在混合自动驾驶交通环境中,每辆自动驾驶汽车的决策都可能影响整个交通系统。由于车辆间复杂的交互关系,如何在保证高通行效率和安全性的前提下做出决策仍具挑战。连接式自动驾驶车辆(CAVs)凭借更强的感知与通信能力,有望显著提升此类动态交互环境中的决策质量。针对基于深度强化学习(DRL)的多车协同决策算法,需有效表征车辆间交互以提取交互特征,而表征方式直接影响学习效率与策略质量。为此,本文提出一种基于Transformer与强化学习的CAV决策架构:引入可学习策略令牌作为多车联合策略的学习媒介,使所有关注区域内的车辆状态能自适应注意,从而提取智能体间的交互特征;同时设计直观的物理位置编码,冗余位置信息优化网络性能。仿真结果表明,本模型能充分融合交通场景中所有车辆的状态信息,生成兼顾效率与安全的高质量驾驶决策,相较现有DRL方法有显著提升。

原文摘要 · Abstract (English)

In mixed autonomy traffic environment, every decision made by an autonomous-driving car may have a great impact on the transportation system. Because of the complex interaction between vehicles, it is challenging to make decisions that can ensure both high traffic efficiency and safety now and futher. Connected automated vehicles (CAVs) have great potential to improve the quality of decision-making in this continuous, highly dynamic and interactive environment because of their stronger sensing and communicating ability. For multi-vehicle collaborative decision-making algorithms based on deep reinforcement learning (DRL), we need to represent the interactions between vehicles to obtain interactive features. The representation in this aspect directly affects the learning efficiency and the quality of the learned policy. To this end, we propose a CAV decision-making architecture based on transformer and reinforcement learning algorithms. A learnable policy token is used as the learning medium of the multi-vehicle joint policy, the states of all vehicles in the area of interest can be adaptively noticed in order to extract interactive features among agents. We also design an intuitive physical positional encodings, the redundant location information of which optimizes the performance of the network. Simulations show that our model can make good use of all the state information of vehicles in traffic scenario, so as to obtain high-quality driving decisions that meet efficiency and safety objectives. The comparison shows that our method significantly improves existing DRL-based multi-vehicle cooperative decision-making algorithms.

自动驾驶强化学习Transformer多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。