arXiv:2512.21831cs.CV2025-12

融合多视角多模态数据,提升自动驾驶中复杂场景下的感知精度。

End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration

  • 通过双层空间交叉注意力实现异构视角与模态对齐。
  • 在真实与模拟数据集上,检测与追踪性能提升15%-20%。
  • 适合需要高鲁棒性感知的自动驾驶系统研发者参考。

多视角协同感知与多模态融合对自动驾驶中的可靠三维时空理解至关重要,尤其在存在遮挡、视角受限及车联万物(V2X)通信延迟的场景下。本文提出一种面向V2X协作的跨模态端到端跟踪框架XET-V2X,该框架在共享时空表示中统一多视角多模态感知。为高效对齐异构视角与模态,XET-V2X引入基于多尺度可变形注意力的双层空间交叉注意力模块。多视角图像特征被聚合以增强语义一致性,随后由更新的空间查询引导点云融合,实现有效的跨模态交互并降低计算开销。基于真实世界V2X序列感知数据集(V2X-Seq-SPD)及两个模拟生成子集(车辆到车辆:V2X-Sim-V2V;车辆到基础设施:V2X-Sim-V2I)的实验表明,在不同通信延迟条件下,该方法持续提升检测与追踪性能,相比单视角或单模态基线,平均精度(mAP)与平均多目标追踪准确率(AMOTA)最高提升15%-20%,同时优于代表性检测后追踪的协同感知方法。

原文摘要 · Abstract (English)

Multiview cooperative perception and multimodal fusion are essential for reliable 3-D spatiotemporal understanding in autonomous driving, especially in cases with occlusions, limited viewpoints, and communication delays in vehicle-to-everything (V2X) scenarios. In this paper, Cross-modal End-to-End Tracking for V2X (XET-V2X), a multimodal fused end-to-end tracking framework for V2X collaboration that unifies multiview multimodal sensing within a shared spatiotemporal representation, is proposed. To efficiently align heterogeneous viewpoints and modalities, XET-V2X introduces a dual-layer spatial cross-attention module based on multiscale deformable attention. Multiview image features are aggregated to enhance semantic consistency, followed by point cloud fusion guided by the updated spatial queries, enabling effective cross-modal interaction while reducing computational overhead. Experiments based on the real-world V2X Sequential Perception Dataset (V2X-Seq-SPD) dataset and two simulated V2X-Sim-derived subsets, namely the vehicle-to-vehicle (V2X-Sim-V2V) and vehicle-to-infrastructure (V2X-Sim-V2I) subsets, demonstrate consistent improvements in detection and tracking performance under varying communication delays, with XET-V2X achieving up to 15-20% relative gains in mean average precision (mAP) and average multi-object tracking accuracy (AMOTA) over single-view or single-modal baselines, while also outperforming representative tracking-by-detection cooperative perception methods.

自动驾驶多模态融合V2X3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。