arXiv:2511.09013cs.ROcs.CV2025-11被引 8

提出多智能体协同框架,提升自动驾驶感知、预测与规划一致性

UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving

  • 通过多层次融合共享查询,实现感知与预测的协同推理
  • 在DAIR-V2X上实现感知准确率提升39.7%,规划性能提高33.2%
  • 引入MoE动态增强特征表示,适合复杂交通场景下的端到端系统

自动驾驶虽具变革潜力,但受限于单体智能的感知局限与孤立决策。现有多智能体方法多聚焦感知层协作,忽视与下游规划控制的对齐,或未能充分发挥端到端系统潜力。本文提出UniMM-V2X,一种端到端多智能体框架,支持感知、预测与规划的分层协同。核心是多层级融合策略,统一感知与预测协作,使智能体共享查询并协同推理,确保决策一致与安全。为适应多样化任务并增强融合质量,引入混合专家(MoE)架构动态优化鸟瞰图(BEV)表示,并将MoE扩展至解码器以捕捉多样运动模式。在DAIR-V2X数据集上的大量实验表明,相比UniV2X,本方法感知准确率提升39.7%,预测误差降低7.2%,规划性能提高33.2%,验证了所提MoE增强型多层级协同范式的优势。

原文摘要 · Abstract (English)

Autonomous driving holds transformative potential but remains fundamentally constrained by the limited perception and isolated decision-making with standalone intelligence. While recent multi-agent approaches introduce cooperation, they often focus merely on perception-level tasks, overlooking the alignment with downstream planning and control, or fall short in leveraging the full capacity of the recent emerging end-to-end autonomous driving. In this paper, we present UniMM-V2X, a novel end-to-end multi-agent framework that enables hierarchical cooperation across perception, prediction, and planning. At the core of our framework is a multi-level fusion strategy that unifies perception and prediction cooperation, allowing agents to share queries and reason cooperatively for consistent and safe decision-making. To adapt to diverse downstream tasks and further enhance the quality of multi-level fusion, we incorporate a Mixture-of-Experts (MoE) architecture to dynamically enhance the BEV representations. We further extend MoE into the decoder to better capture diverse motion patterns. Extensive experiments on the DAIR-V2X dataset demonstrate our approach achieves state-of-the-art (SOTA) performance with a 39.7% improvement in perception accuracy, a 7.2% reduction in prediction error, and a 33.2% improvement in planning performance compared with UniV2X, showcasing the strength of our MoE-enhanced multi-level cooperative paradigm.

自动驾驶多智能体端到端MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。