提出V2XPnP框架,实现多智能体时空协同感知与预测
V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction
- 设计一跳通信下三种融合策略的完整对比实验
- 在真实数据集上实现感知与预测双任务性能领先
- 支持全模式协作的序列化数据集,填补现有空白
车联网(V2X)技术为缓解单车观测受限问题提供了新范式。现有工作多聚焦单帧协同感知,忽视时间线索及时间任务(如时序感知与预测)。本文聚焦V2X场景中的时空融合,设计一跳与多跳通信策略(何时传),并考察其与三种融合策略——早期、晚期、中间融合(传什么)的集成,提供包含11种融合模型的全面基准。进一步提出V2XPnP,一种基于统一Transformer架构的中间融合框架,适用于一跳通信下的端到端感知与预测。该框架能有效建模多智能体、多帧及高精地图间的复杂时空关系。同时构建V2XPnP顺序数据集,支持所有V2X协作模式,解决现有真实数据集仅限单帧或单模式协作的局限。大量实验证明,该框架在感知与预测任务上均优于现有最优方法。
原文摘要 · Abstract (English)
Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。