arXiv:2509.09946cs.CV2025-09ICCV被引 3

用深度信息将2D跟踪扩展到3D,无需重写系统即可实现多摄像头三维目标追踪。

Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation

  • 基于深度图重建目标点云,通过聚类和航向优化恢复3D框
  • 在AI City Challenge 3D数据集上获第三名,保持在线推理能力
  • 适用于已有2D跟踪系统的快速3D升级,适合智能监控场景

多目标多摄像头跟踪(MTMC)是自动化大规模监控的关键计算机视觉任务。结合相机标定与深度信息,可将场景中的目标投影至3D空间,实现对3D环境的高精度自动感知。然而,直接在3D空间中进行跟踪需重构全部2D跟踪组件,对现有系统不友好。本文提出一种方法,通过深度信息将任意在线2D多摄像头跟踪系统扩展至3D空间:先在点云空间重建目标,再经聚类与航向修正恢复3D边界框。同时引入改进的在线数据关联机制,利用目标局部ID一致性分配全局ID。该框架在2025 AI City Challenge的3D MTMC数据集上表现优异,位列排行榜第3名。

原文摘要 · Abstract (English)

Multi-Target Multi-Camera Tracking (MTMC) is an essential computer vision task for automating large-scale surveillance. With camera calibration and depth information, the targets in the scene can be projected into 3D space, offering unparalleled levels of automatic perception of a 3D environment. However, tracking in the 3D space requires replacing all 2D tracking components from the ground up, which may be infeasible for existing MTMC systems. In this paper, we present an approach for extending any online 2D multi-camera tracking system into 3D space by utilizing depth information to reconstruct a target in point-cloud space, and recovering its 3D box through clustering and yaw refinement following tracking. We also introduced an enhanced online data association mechanism that leverages the target's local ID consistency to assign global IDs across frames. The proposed framework is evaluated on the 2025 AI City Challenge's 3D MTMC dataset, achieving 3rd place on the leaderboard.

多摄像头跟踪3D感知深度信息在线追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。