arXiv:2608.07621cs.AIcs.CV2026-08

提出首个闭环协作自动驾驶基准与模型,支持多车协同决策。

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

论文配图:CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建多车协同驾驶的闭环评估基准CMU-Drive。
  • 提出V2V-VLA模型,单次前向传播完成动作、路径、语言推理与通信生成。
  • 适合研究多智能体协同自动驾驶的学者与工程师。

视觉-语言-动作(VLA)模型在端到端自动驾驶中表现优异,但现有方法主要针对单一自动驾驶代理,缺乏对协作感知、推理和规划的支持。本文提出协作多智能体统一驾驶推理基准CMU-Drive,用于评估多个联网自动驾驶车辆(CAVs)在包含背景交通参与者的高危驾驶场景中的表现。同时提出车对车视觉-语言-动作(V2V-VLA)模型,通过联合生成驾驶动作、未来航点、语言推理与通信策略,在一次前向传播中实现协作驾驶。在CMU-Drive上的实验建立了首个协作式VLA驾驶基准与基线,为未来多智能体、闭环、端到端协同自动驾驶研究奠定基础。代码、基准与模型检查点将公开发布,促进开源研究。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.

自动驾驶多智能体视觉语言动作协同决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。