arXiv:2606.21165cs.ROcs.AI2026-06中稿 · IROS 2026

OmniV2X用生成模型实现高效协同驾驶,降低通信与计算开销。

OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving

论文配图:OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving
图 1 · 摘自论文原文
  • 通过跨注意力注入,直接处理多模态多智能体输入,避免共享表示瓶颈。
  • 仅用1%通信带宽和不到10%标注数据,达到顶尖协同驾驶性能。
  • 适合自动驾驶系统研发,尤其关注低延迟与高鲁棒性的场景。

我们提出OmniV2X,一种面向车联万物(V2X)协同驾驶的生成基础模型。该模型直接解析包含多模态与多智能体观测的独立上下文序列,解决了密集3D感知带来的高计算成本、协同场景下数据稀缺的脆弱性以及现有方法融合多模态输入时对标准消息格式的不兼容问题。训练采用端到端监督流程,以轨迹生成损失为目标,高容量生成序列规划器通过交叉注意力注入隐式学习如何引导模型并利用多模态输入。作为基础模型,OmniV2X在大规模单智能体规划数据集上预训练后,可通过轻量级、符合标准的V2X token有效适配协同环境。在DAIR-V2X-Seq数据集上的评估表明,其性能优于现有端到端协同驾驶基线,仅需不到10%的微调V2X数据集和不到1%的通信带宽即可达成最先进效果。全面评估验证了其在真实约束下的计算效率与鲁棒性。

原文摘要 · Abstract (English)

We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets independent context sequences comprising multi-modal and multi-agent observations. The new design mitigates the computational cost of dense 3D perception, the vulnerability to data scarcity in cooperative scenarios, and the poor compliance with standardized messaging in existing methods that fuse multi-modal inputs into a shared representation. For training, we present an end-to-end supervised pipeline using a downstream trajectory generation loss, in which a high-capacity generative sequence planner implicitly learns to steer the model and leverage multi-modal inputs via cross-attention injection. As a foundation model, we demonstrate that OmniV2X pre-trained on large-scale single-agent planning datasets can efficiently adapt to cooperative environments by integrating the conditioning context with lightweight, standard-compliant V2X tokens. Evaluated on the DAIR-V2X-Seq dataset, OmniV2X outperforms existing end-to-end cooperative driving baselines, achieving state-of-the-art performance with less than 10% of the fine-tune V2X dataset and less than 1% of the communication bandwidth. We conduct comprehensive evaluations to demonstrate its computational efficiency and robustness under real-world constraints.

协同驾驶生成模型V2X端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。