用重叠分块提升耳部识别准确率,效果显著。
Improved Ear Verification with Vision Transformers and Overlapping Patches
- 采用重叠分块策略增强ViT对耳部细节的捕捉能力。
- 在48组实验中44组表现更优,最高提升10%准确率。
- 小模型ViT-T在多个数据集上表现最佳,适合实际部署。
耳部识别因成年后外观稳定而成为有前景的生物特征识别方式。尽管视觉变换器(ViTs)广泛应用于图像识别,但其在耳部识别中的效率受限于对重叠分块的忽视,而重叠分块对捕捉复杂耳部特征至关重要。本研究在多样化的数据集(OPIB、AWE、WPUT和EarVN1.0)上评估了ViT-Tiny(ViT-T)、ViT-Small(ViT-S)、ViT-Base(ViT-B)和ViT-Large(ViT-L)配置,采用重叠分块策略。结果表明,重叠分块极为重要,在48组结构化实验中有44组表现更优。与非重叠配置相比,性能提升显著,最高达10%(在EarVN1.0数据集上)。在模型表现上,ViT-T在AWE、WPUT和EarVN1.0数据集上持续优于其他模型。最佳配置为28×28分块大小、14像素步幅,对应归一化图像区域的25%(112×112像素)和行列尺寸的12.5%。该研究证实,结合重叠分块的变压器架构可作为耳部生物特征验证任务中高效且高性能的选择。
原文摘要 · Abstract (English)
Ear recognition has emerged as a promising biometric modality due to the relative stability in appearance during adulthood. Although Vision Transformers (ViTs) have been widely used in image recognition tasks, their efficiency in ear recognition has been hampered by a lack of attention to overlapping patches, which is crucial for capturing intricate ear features. In this study, we evaluate ViT-Tiny (ViT-T), ViT-Small (ViT-S), ViT-Base (ViT-B) and ViT-Large (ViT-L) configurations on a diverse set of datasets (OPIB, AWE, WPUT, and EarVN1.0), using an overlapping patch selection strategy. Results demonstrate the critical importance of overlapping patches, yielding superior performance in 44 of 48 experiments in a structured study. Moreover, upon comparing the results of the overlapping patches with the non-overlapping configurations, the increase is significant, reaching up to 10% for the EarVN1.0 dataset. In terms of model performance, the ViT-T model consistently outperformed the ViT-S, ViT-B, and ViT-L models on the AWE, WPUT, and EarVN1.0 datasets. The highest scores were achieved in a configuration with a patch size of 28x28 and a stride of 14 pixels. This patch-stride configuration represents 25% of the normalized image area (112x112 pixels) for the patch size and 12.5% of the row or column size for the stride. This study confirms that transformer architectures with overlapping patch selection can serve as an efficient and high-performing option for ear-based biometric recognition tasks in verification scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。