arXiv:2504.03047cs.CV2025-04被引 1

用注意力机制提升多视角行人追踪的鲁棒性,有效缓解透视畸变影响。

Attention-Aware Multi-View Pedestrian Tracking

  • 多视角特征早期融合至鸟瞰图,结合跨帧注意力关联行人
  • 在Wildtrack上达到96.1% IDF1,MultiviewX上达85.7%
  • 适合需要高精度多摄像头行人追踪的场景

尽管多目标追踪技术取得进展,遮挡仍是主要挑战。多相机系统通过提供全景覆盖缓解该问题。近期多视角行人检测模型表明,将各视角特征图投影至统一地面平面或鸟瞰图(BEV)进行早期融合可提升检测与追踪性能。然而,透视变换导致地面平面出现显著畸变,影响行人外观特征的鲁棒性。为此,本文提出一种新型多视角行人追踪模型,在检测阶段采用早期融合策略,并引入跨注意力机制,在不同帧间建立稳健的行人关联,同时高效传播行人特征,从而获得更鲁棒的特征表示。大量实验表明,本模型优于当前最优方法:在Wildtrack数据集上取得96.1%的IDF1,在MultiviewX数据集上达到85.7%。

原文摘要 · Abstract (English)

In spite of the recent advancements in multi-object tracking, occlusion poses a significant challenge. Multi-camera setups have been used to address this challenge by providing a comprehensive coverage of the scene. Recent multi-view pedestrian detection models have highlighted the potential of an early-fusion strategy, projecting feature maps of all views to a common ground plane or the Bird's Eye View (BEV), and then performing detection. This strategy has been shown to improve both detection and tracking performance. However, the perspective transformation results in significant distortion on the ground plane, affecting the robustness of the appearance features of the pedestrians. To tackle this limitation, we propose a novel model that incorporates attention mechanisms in a multi-view pedestrian tracking scenario. Our model utilizes an early-fusion strategy for detection, and a cross-attention mechanism to establish robust associations between pedestrians in different frames, while efficiently propagating pedestrian features across frames, resulting in a more robust feature representation for each pedestrian. Extensive experiments demonstrate that our model outperforms state-of-the-art models, with an IDF1 score of $96.1\%$ on Wildtrack dataset, and $85.7\%$ on MultiviewX dataset.

行人追踪多视角注意力机制鸟瞰图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。