提出COIN框架,让自动驾驶车在复杂路况中更安全高效协同
COIN: Collaborative Interaction-Aware Multi-Agent Reinforcement Learning for Self-Driving Systems
- 采用联合优化个体与全局目标的双层交互感知算法
- 在密集城市交通中提升安全与效率,优于现有方法
- 适合研究智能交通系统协同决策的学者与工程师
多智能体自动驾驶(MASD)系统为协调自动驾驶车辆、减少拥堵并提升未来智能交通系统的安全性和运行效率提供了有效方案。多智能体强化学习(MARL)已成为构建先进端到端MASD系统的重要途径。然而,在动态场景下实现高效且安全的协作仍是密集环境中复杂交互下的重大挑战。为此,本文提出一种新型协同(CO-)交互感知(-IN)MARL框架——COIN。具体而言,我们设计了一种新的反事实个体-全局孪生延迟深度确定性策略梯度(CIG-TD3)算法,采用“集中训练、分散执行”(CTDE)模式,旨在联合优化智能体的个体目标(导航)与全局目标(协作)。我们进一步引入双层交互感知的集中式评判器架构,捕捉局部成对交互与全局系统级依赖关系,从而实现更准确的全局价值估计和改进的信用分配,促进协作策略学习。我们在密集城市交通环境中的大量仿真实验表明,COIN在不同系统规模下均持续优于其他先进基线方法,在安全性和效率方面表现更优。结果凸显其在复杂动态MASD场景中的优越性,通过真实世界机器人演示进一步验证。补充视频见https://marmotlab.github.io/COIN/
原文摘要 · Abstract (English)
Multi-Agent Self-Driving (MASD) systems provide an effective solution for coordinating autonomous vehicles to reduce congestion and enhance both safety and operational efficiency in future intelligent transportation systems. Multi-Agent Reinforcement Learning (MARL) has emerged as a promising approach for developing advanced end-to-end MASD systems. However, achieving efficient and safe collaboration in dynamic MASD systems remains a significant challenge in dense scenarios with complex agent interactions. To address this challenge, we propose a novel collaborative(CO-) interaction-aware(-IN) MARL framework, named COIN. Specifically, we develop a new counterfactual individual-global twin delayed deep deterministic policy gradient (CIG-TD3) algorithm, crafted in a "centralized training, decentralized execution" (CTDE) manner, which aims to jointly optimize the individual objectives (navigation) and the global objectives (collaboration) of agents. We further introduce a dual-level interaction-aware centralized critic architecture that captures both local pairwise interactions and global system-level dependencies, enabling more accurate global value estimation and improved credit assignment for collaborative policy learning. We conduct extensive simulation experiments in dense urban traffic environments, which demonstrate that COIN consistently outperforms other advanced baseline methods in both safety and efficiency across various system sizes. These results highlight its superiority in complex and dynamic MASD scenarios, as further validated through real-world robot demonstrations. Supplementary videos are available at https://marmotlab.github.io/COIN/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。