arXiv:2507.19239cs.CV2025-07ICCV被引 19

提出端到端协作跟踪框架,显著提升多车协同感知效率

CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception

  • 全实例级端到端设计,可学习实例关联机制
  • 在V2X-Seq上达39.0% mAP、32.8% AMOTA,性能领先
  • 适合研究多智能体协同感知与车载系统优化的开发者

协作感知通过多智能体间信息共享,解决单车自动驾驶系统的固有局限。以往研究多集中于单帧感知任务,而更复杂的协作序列感知任务(如协作3D多目标跟踪)尚未充分探索。为此,本文提出CoopTrack,一种全实例级端到端协作跟踪框架,具备可学习的实例关联能力,与现有方法本质不同。该框架传输稀疏实例级特征,在显著提升感知能力的同时保持低通信开销。其核心包含两个组件:多维特征提取,以及跨智能体关联与聚合,共同实现融合语义与运动信息的完整实例表征,并基于特征图实现自适应跨智能体关联与融合。在V2X-Seq和Griffin数据集上的实验表明,CoopTrack表现优异,尤其在V2X-Seq上达到39.0% mAP和32.8% AMOTA的先进水平。项目代码已开源。

原文摘要 · Abstract (English)

Cooperative perception aims to address the inherent limitations of single-vehicle autonomous driving systems through information exchange among multiple agents. Previous research has primarily focused on single-frame perception tasks. However, the more challenging cooperative sequential perception tasks, such as cooperative 3D multi-object tracking, have not been thoroughly investigated. Therefore, we propose CoopTrack, a fully instance-level end-to-end framework for cooperative tracking, featuring learnable instance association, which fundamentally differs from existing approaches. CoopTrack transmits sparse instance-level features that significantly enhance perception capabilities while maintaining low transmission costs. Furthermore, the framework comprises two key components: Multi-Dimensional Feature Extraction, and Cross-Agent Association and Aggregation, which collectively enable comprehensive instance representation with semantic and motion features, and adaptive cross-agent association and fusion based on a feature graph. Experiments on both the V2X-Seq and Griffin datasets demonstrate that CoopTrack achieves excellent performance. Specifically, it attains state-of-the-art results on V2X-Seq, with 39.0\% mAP and 32.8\% AMOTA. The project is available at https://github.com/zhongjiaru/CoopTrack.

协同感知3D跟踪端到端车联网

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。