arXiv:2412.00433cs.CV2024-12被引 7

动态选关键图像片段,提升空地人物识别准确率

Dynamic Token Selection for Aerial-Ground Person Re-Identification

  • 按重要性动态筛选图像关键区域,聚焦身份特征
  • 在CARGO数据集上比第二名高1.18% mAP
  • 适合处理视角、光照差异大的空地人物识别场景

空地人物再识别(AGPReID)具有重要实用价值,但因视角、光照和背景干扰差异显著而面临挑战。传统方法常全局分析整图,易受无关信息影响且效率低。本文提出专用于AGPReID的动态令牌选择变压器(DTST),通过将输入图像分割为多个令牌(每令牌代表一个区域或特征),采用Top-k策略选取最相关的k个关键令牌,集中关注对身份识别至关重要的区域。随后利用注意力机制挖掘不同令牌间的关联,增强身份特征表示。在基准数据集上的大量实验表明,本方法优于现有工作。特别地,在CARGO数据集上,相比第二名提升1.18% mAP。此外,我们系统分析了令牌数量、插入位置及注意力头数对模型性能的影响。

原文摘要 · Abstract (English)

Aerial-Ground Person Re-identification (AGPReID) holds significant practical value but faces unique challenges due to pronounced variations in viewing angles, lighting conditions, and background interference. Traditional methods, often involving a global analysis of the entire image, frequently lead to inefficiencies and susceptibility to irrelevant data. In this paper, we propose a novel Dynamic Token Selective Transformer (DTST) tailored for AGPReID, which dynamically selects pivotal tokens to concentrate on pertinent regions. Specifically, we segment the input image into multiple tokens, with each token representing a unique region or feature within the image. Using a Top-k strategy, we extract the k most significant tokens that contain vital information essential for identity recognition. Subsequently, an attention mechanism is employed to discern interrelations among diverse tokens, thereby enhancing the representation of identity features. Extensive experiments on benchmark datasets showcases the superiority of our method over existing works. Notably, on the CARGO dataset, our proposed method gains 1.18% mAP improvements when compared to the second place. In addition, we comprehensively analyze the impact of different numbers of tokens, token insertion positions, and numbers of heads on model performance.

人物重识别动态选择视觉Transformer空地协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。