用紧凑嵌入统一协调无人车过无灯路口,效果更好更通用。
Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections

- 中心主代理生成全局协调嵌入,分发给各车辆执行。
- 72种路口配置下零碰撞,平均通行时间显著降低。
- 三车训练可直接用于五车场景,泛化能力强。
在无灯交叉口协调自动驾驶车辆仍是多智能体强化学习(MARL)的核心挑战,现有方法常受限于组合动作空间、依赖特权信息或僵化的代理设计。本文提出主代理原型计划系统(MAPS),一种分层深度强化学习架构:中心主代理生成一个紧凑的连续嵌入(称为原型计划),编码全局协调策略;去中心化的工作者代理将该嵌入与本地观测融合,执行车辆特定控制,从而解耦战略意图与战术执行,实现模块独立优化。作为该协调机制的验证,我们在HighwayEnv中测试了72种交叉口配置。MAPS实现了零碰撞导航,同时显著降低平均通行时间,优于现有最优基线。学习到的原型计划还展现出强泛化能力:在三车场景训练后,零样本部署至五车场景仍达到94%成功率,证实基于原型计划的分层学习为多车协调提供了有前景的框架。
原文摘要 · Abstract (English)
Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs. We propose Master-Agent Proto-plan System (MAPS), a hierarchical deep reinforcement learning (DRL) architecture in which a centralized Master agent generates a compact, continuous embedding, denoted as proto-plan, that encodes a global coordination strategy. Decentralized Worker agents integrate this embedding with local observations to execute vehicle-specific control, decoupling strategic intent from tactical execution and enabling independent optimization of each module. As a proof-of-concept evaluation of this coordination mechanism, we test MAPS across 72 intersection configurations in HighwayEnv. MAPS achieves collision-free navigation while significantly reducing average travel time, outperforming state-of-the-art baselines. The learned proto-plans further exhibit robust generalization: a system trained with three agents achieves a 94% success rate when deployed zero-shot to five-agent scenarios, confirming that proto-plan-based hierarchical learning provides a promising framework for multi-vehicle coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。