多视角视频中3D目标检测与跟踪,提升长时关联能力
MCBLT: Multi-Camera Multi-Object 3D Tracking in Long Videos
- 通过多视角图像融合生成鸟瞰图3D检测结果
- 在AICity'24上达81.22 HOTA,WildTrack上95.6 IDF1
- 适用于不同场景和相机布局,适合长视频追踪
多视角摄像头中的目标感知对智能系统至关重要,尤其在仓库、零售店和医院等室内环境中。现有主流多目标多相机(MTMC)检测与跟踪方法依赖2D检测、单视图多目标跟踪(MOT)和跨视角重识别(ReID),未能有效利用多视角图像聚合带来的3D信息。本文提出一种3D目标检测与跟踪框架MCBLT,首先利用相机标定参数融合多视角图像,生成鸟瞰图(BEV)中的3D目标检测结果;随后引入分层图神经网络(GNN)在BEV空间中追踪这些3D检测,实现MTMC跟踪。相较于现有方法,MCBLT具备出色的跨场景泛化能力,且在长时关联任务中表现优异。在AICity'24数据集上取得81.22 HOTA,在WildTrack数据集上达到95.6 IDF1,建立新基准。
原文摘要 · Abstract (English)
Object perception from multi-view cameras is crucial for intelligent systems, particularly in indoor environments, e.g., warehouses, retail stores, and hospitals. Most traditional multi-target multi-camera (MTMC) detection and tracking methods rely on 2D object detection, single-view multi-object tracking (MOT), and cross-view re-identification (ReID) techniques, without properly handling important 3D information by multi-view image aggregation. In this paper, we propose a 3D object detection and tracking framework, named MCBLT, which first aggregates multi-view images with necessary camera calibration parameters to obtain 3D object detections in bird's-eye view (BEV). Then, we introduce hierarchical graph neural networks (GNNs) to track these 3D detections in BEV for MTMC tracking results. Unlike existing methods, MCBLT has impressive generalizability across different scenes and diverse camera settings, with exceptional capability for long-term association handling. As a result, our proposed MCBLT establishes a new state-of-the-art on the AICity'24 dataset with $81.22$ HOTA, and on the WildTrack dataset with $95.6$ IDF1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。