用视觉大模型提升多类行人车辆协同检测跟踪精度
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
- 引入全局空间注意力融合模块增强多尺度特征学习
- 基于视觉大模型的重识别模块降低小物体误追踪率
- 动态调整跟踪间隔,适应不同运动速度的目标
协同感知通过扩展感知范围和提升对传感器故障的鲁棒性,显著增强环境理解能力,主要涉及协同3D检测与跟踪任务。前者关注单帧中的目标识别,后者则追踪跨时间的实例轨迹。然而,现有方法多集中于车辆类别,缺乏对多类目标(如行人、非机动车)协同检测与跟踪的有效解决方案,限制了其在多样化真实场景中的应用。为此,本文提出面向多样道路使用者的多类协同检测与跟踪框架。首先设计具备全局空间注意力融合(GSAF)模块的检测器,强化不同尺寸目标的多尺度特征学习;其次引入基于视觉基础模型的轨迹重识别(REID)模块,利用视觉语义信息有效减少小目标(如行人)的ID切换错误;最后提出基于速度的自适应轨迹管理(VATM)模块,根据物体运动状态动态调整跟踪间隔。在V2X-Real和OPV2V数据集上的大量实验表明,该方法在检测与跟踪准确率上均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Collaborative perception plays a crucial role in enhancing environmental understanding by expanding the perceptual range and improving robustness against sensor failures, which primarily involves collaborative 3D detection and tracking tasks. The former focuses on object recognition in individual frames, while the latter captures continuous instance tracklets over time. However, existing works in both areas predominantly focus on the vehicle superclass, lacking effective solutions for both multi-class collaborative detection and tracking. This limitation hinders their applicability in real-world scenarios, which involve diverse object classes with varying appearances and motion patterns. To overcome these limitations, we propose a multi-class collaborative detection and tracking framework tailored for diverse road users. We first present a detector with a global spatial attention fusion (GSAF) module, enhancing multi-scale feature learning for objects of varying sizes. Next, we introduce a tracklet RE-IDentification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID SWitch (IDSW) errors, in cases of erroneous mismatches involving small objects like pedestrians. We further design a velocity-based adaptive tracklet management (VATM) module that adjusts the tracking interval dynamically based on object motion. Extensive experiments on the V2X-Real and OPV2V datasets show that our approach significantly outperforms existing state-of-the-art methods in both detection and tracking accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。