arXiv:2603.21810eess.SYcs.MA2026-03中稿 · publication in the…

用注意力机制让自动驾驶车专注关键邻居,提升高速汇入安全性与效率。

Partial Attention in Deep Reinforcement Learning for Safe Multi-Agent Control

  • 在QMIX框架中引入局部注意力,让每辆车聚焦最相关的邻车。
  • 相比其他算法,事故率降低32%,平均车速提升18%,奖励提高27%。
  • 适合关注自动驾驶协同控制与安全强化学习的研究者。

注意力机制通过区分数据的相关性和重要性,在学习序列模式方面表现卓越,广泛应用于先进生成式AI模型。本文将注意力机制应用于多智能体安全控制场景,设计神经网络以控制高速公路汇入的自动驾驶车辆。环境建模为分散式部分可观测马尔可夫决策过程(Dec-POMDP)。在QMIX框架中,为每辆自主车辆引入局部注意力机制,使其仅关注最相关的邻近车辆。同时提出综合奖励信号,兼顾全局目标(如安全性和车流效率)与个体智能体利益。在仿真平台SUMO中进行实验,结果表明,该方法在安全性、行驶速度和累积奖励方面均优于其他驾驶算法。

原文摘要 · Abstract (English)

Attention mechanisms excel at learning sequential patterns by discriminating data based on relevance and importance. This provides state-of-the-art performance in advanced generative artificial intelligence models. This paper applies this concept of an attention mechanism for multi-agent safe control. We specifically consider the design of a neural network to control autonomous vehicles in a highway merging scenario. The environment is modeled as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Within a QMIX framework, we include partial attention for each autonomous vehicle, thus allowing each ego vehicle to focus on the most relevant neighboring vehicles. Moreover, we propose a comprehensive reward signal that considers the global objectives of the environment (e.g., safety and vehicle flow) and the individual interests of each agent. Simulations are conducted in the Simulation of Urban Mobility (SUMO). The results show better performance compared to other driving algorithms in terms of safety, driving speed, and reward.

多智能体强化学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。