arXiv:2411.06297cs.CV2024-11被引 5

针对车辆重识别中非正方形图像影响精度的问题,提出自适应长宽比的ViT框架。

Adaptive Aspect Ratios with Patch-Mixup-ViT-based Vehicle ReID

  • 基于空间注意力引导的块混合策略,适配物体真实长宽比
  • 在VeRi-776和VehicleID上超越现有SOTA方法,推理时间几乎不变
  • 适合处理长宽比多变的真实场景车辆图像,提升模型鲁棒性

视觉变换器(ViTs)在车辆重识别(ReID)任务中表现优异,但输入图像或视频的非正方形长宽比会降低识别准确率。为此,本文提出一种受人类感知启发、通用性强的基于ViT的ReID框架,融合在多种长宽比下训练的模型。主要贡献包括:(i) 利用VeRi-776和VehicleID数据集分析长宽比对性能的影响,为输入设置提供依据;(ii) 在ViT的图像分块阶段引入基于空间注意力得分的块级混合策略,并采用不均匀步长以更好匹配目标物体的长宽比;(iii) 设计动态特征融合网络增强模型鲁棒性。该方法在两个数据集上均优于现有SOTA Transformer方法,且每张图像推理时间仅小幅增加。

原文摘要 · Abstract (English)

Vision Transformers (ViTs) have shown exceptional performance in vehicle re-identification (ReID) tasks. However, non-square aspect ratios of image or video inputs can negatively impact re-identification accuracy. To address this challenge, we propose a novel, human perception driven, and general ViT-based ReID framework that fuses models trained on various aspect ratios. Our key contributions are threefold: (i) We analyze the impact of aspect ratios on performance using the VeRi-776 and VehicleID datasets, providing guidance for input settings based on the distribution of original image aspect ratios. (ii) We introduce patch-wise mixup strategy during ViT patchification (guided by spatial attention scores) and implement uneven stride for better alignment with object aspect ratios. (iii) We propose a dynamic feature fusion ReID network to enhance model robustness. Our method outperforms state-of-the-art transformer-based approaches on both datasets, with only a minimal increase in inference time per image.

车辆重识别ViT长宽比自适应特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。