arXiv:2501.10129cs.CVcs.AI2025-01被引 1

通过自适应关键帧挖掘与帧内特征融合,提升多目标追踪在遮挡下的准确率。

Spatio-temporal Graph Learning on Adaptive Mined Key Frames for High-performance Multi-Object Tracking

  • 用强化学习自适应提取关键帧,更好捕捉视频内在时空关系。
  • 在MOT17上达68.6 HOTA、81.0 IDF1,误跟踪数仅893。
  • 适合需要高精度追踪复杂遮挡场景的研究者与应用开发者。

在多目标追踪领域,准确捕捉视频序列中物体间的时空关系仍是重大挑战,尤其频繁的相互遮挡易引发追踪错误并降低性能。为此,本文提出一种自适应关键帧挖掘策略,设计关键帧提取(KFE)模块,利用强化学习自适应分割视频,引导追踪器挖掘视频内容的内在逻辑。该方法有效捕获物体间的结构化空间关系及跨帧时间关系。为解决遮挡问题,引入帧内特征融合(IFF)模块,采用图卷积网络(GCN)实现帧内目标与邻近物体间的信息交换,增强目标可区分性,缓解因遮挡导致的追踪丢失与外观相似性问题。结合长短期轨迹与物体空间关系,所提追踪器在MOT17数据集上取得68.6 HOTA、81.0 IDF1、66.6 AssA和893 IDS的优异表现,验证了其有效性与准确性。

原文摘要 · Abstract (English)

In the realm of multi-object tracking, the challenge of accurately capturing the spatial and temporal relationships between objects in video sequences remains a significant hurdle. This is further complicated by frequent occurrences of mutual occlusions among objects, which can lead to tracking errors and reduced performance in existing methods. Motivated by these challenges, we propose a novel adaptive key frame mining strategy that addresses the limitations of current tracking approaches. Specifically, we introduce a Key Frame Extraction (KFE) module that leverages reinforcement learning to adaptively segment videos, thereby guiding the tracker to exploit the intrinsic logic of the video content. This approach allows us to capture structured spatial relationships between different objects as well as the temporal relationships of objects across frames. To tackle the issue of object occlusions, we have developed an Intra-Frame Feature Fusion (IFF) module. Unlike traditional graph-based methods that primarily focus on inter-frame feature fusion, our IFF module uses a Graph Convolutional Network (GCN) to facilitate information exchange between the target and surrounding objects within a frame. This innovation significantly enhances target distinguishability and mitigates tracking loss and appearance similarity due to occlusions. By combining the strengths of both long and short trajectories and considering the spatial relationships between objects, our proposed tracker achieves impressive results on the MOT17 dataset, i.e., 68.6 HOTA, 81.0 IDF1, 66.6 AssA, and 893 IDS, proving its effectiveness and accuracy.

多目标追踪图神经网络关键帧提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。