arXiv:2502.00665cs.CV2025-02

用静态图中的运动线索提升行人重识别,尤其在遮挡时更准

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios

  • 从静态图像中挖掘运动线索,融合视频特征增强识别
  • 在多类遮挡和视频场景下,超越现有方法表现
  • 适合做复杂监控场景下的行人追踪研究者

在不同监控场景下,尤其是存在遮挡时,行人重识别(ReID)面临巨大挑战。本文提出一种新型的运动感知融合网络(MOTAR-FUSE),利用从静态图像中提取的运动线索显著提升ReID性能。该网络采用双输入视觉适配器,可同时处理图像与视频数据,实现更有效的特征提取。其独特之处在于引入运动一致性任务,使运动感知变压器能够精准捕捉人体运动动态。该方法在遮挡频繁的场景中显著提升特征识别能力,推动ReID进程。我们在多个ReID基准上进行了全面评估,涵盖整体、遮挡及视频场景,结果表明MOTAR-FUSE优于现有方法。

原文摘要 · Abstract (English)

Navigating the complexities of person re-identification (ReID) in varied surveillance scenarios, particularly when occlusions occur, poses significant challenges. We introduce an innovative Motion-Aware Fusion (MOTAR-FUSE) network that utilizes motion cues derived from static imagery to significantly enhance ReID capabilities. This network incorporates a dual-input visual adapter capable of processing both images and videos, thereby facilitating more effective feature extraction. A unique aspect of our approach is the integration of a motion consistency task, which empowers the motion-aware transformer to adeptly capture the dynamics of human motion. This technique substantially improves the recognition of features in scenarios where occlusions are prevalent, thereby advancing the ReID process. Our comprehensive evaluations across multiple ReID benchmarks, including holistic, occluded, and video-based scenarios, demonstrate that our MOTAR-FUSE network achieves superior performance compared to existing approaches.

行人重识别运动感知遮挡处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。