arXiv:2508.08117cs.CVcs.AI2025-08被引 1

通过3D几何推理提升单目视频多目标追踪鲁棒性

GRASPTrack: Geometry-Reasoned Association via Segmentation and Projection for Multi-Object Tracking

  • 融合深度估计与分割生成3D点云,实现精确3D空间关联
  • 在MOT17、MOT20等数据集上显著提升遮挡场景下的追踪精度
  • 适合需要处理复杂运动和频繁遮挡的视觉系统开发者

单目视频中的多目标追踪(MOT)受遮挡和深度模糊的根本挑战,传统检测后追踪(TBD)方法因缺乏几何感知而难以应对。为此,我们提出GRASPTrack,一种将单目深度估计与实例分割融入标准TBD流程的深度感知框架,从2D检测生成高保真3D点云,支持显式3D几何推理。这些点云被体素化,用于精确且鲁棒的体素化3D交并比(IoU)计算。为增强追踪鲁棒性,方法引入深度感知自适应噪声补偿,根据遮挡严重程度动态调整卡尔曼滤波过程噪声。此外,提出深度增强型观测中心动量,将运动方向一致性从图像平面扩展至3D空间,改善复杂轨迹下的运动关联。在MOT17、MOT20和DanceTrack基准上的实验表明,该方法在频繁遮挡与复杂运动场景中显著提升追踪鲁棒性。

原文摘要 · Abstract (English)

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To address these limitations, we introduce GRASPTrack, a novel depth-aware MOT framework that integrates monocular depth estimation and instance segmentation into a standard TBD pipeline to generate high-fidelity 3D point clouds from 2D detections, thereby enabling explicit 3D geometric reasoning. These 3D point clouds are then voxelized to enable a precise and robust Voxel-Based 3D Intersection-over-Union (IoU) for spatial association. To further enhance tracking robustness, our approach incorporates Depth-aware Adaptive Noise Compensation, which dynamically adjusts the Kalman filter process noise based on occlusion severity for more reliable state estimation. Additionally, we propose a Depth-enhanced Observation-Centric Momentum, which extends the motion direction consistency from the image plane into 3D space to improve motion-based association cues, particularly for objects with complex trajectories. Extensive experiments on the MOT17, MOT20, and DanceTrack benchmarks demonstrate that our method achieves competitive performance, significantly improving tracking robustness in complex scenes with frequent occlusions and intricate motion patterns.

多目标追踪3D几何单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。