通过贴合耳部形状的分块变形,提升ViT在耳识别中的鲁棒性。
PaW-ViT: A Patch-based Warping Vision Transformer for Robust Ear Verification
- 基于解剖学知识,将图像分块对齐耳部特征边界
- 在多种ViT模型上实现对形变、尺寸、姿态的稳定识别
- 适合需要高鲁棒性的生物特征认证场景
视觉变换器(ViT)常用矩形令牌进行视觉识别,但其会引入目标外信息,影响性能。本文提出PaW-ViT(基于分块变形的视觉变换器),一种基于解剖学知识的预处理方法,通过将令牌边界精确对齐检测到的耳部特征边界,增强ViT在耳识别中的表现。该方法利用耳部自然曲率对齐特征边界,生成更一致的令牌表示,有效提升对形状、大小和姿态变化的鲁棒性。实验表明,PaW-ViT在多种ViT模型(ViT-T、ViT-S、ViT-B、ViT-L)上均表现出色,具备良好的形变适应能力。本工作旨在解决耳生物特征形态差异与变压器架构位置敏感性之间的脱节问题,为认证系统提供新路径。
原文摘要 · Abstract (English)
The rectangular tokens common to vision transformer methods for visual recognition can strongly affect performance of these methods due to incorporation of information outside the objects to be recognized. This paper introduces PaW-ViT, Patch-based Warping Vision Transformer, a preprocessing approach rooted in anatomical knowledge that normalizes ear images to enhance the efficacy of ViT. By accurately aligning token boundaries to detected ear feature boundaries, PaW-ViT obtains greater robustness to shape, size, and pose variation. By aligning feature boundaries to natural ear curvature, it produces more consistent token representations for various morphologies. Experiments confirm the effectiveness of PaW-ViT on various ViT models (ViT-T, ViT-S, ViT-B, ViT-L) and yield reasonable alignment robustness to variation in shape, size, and pose. Our work aims to solve the disconnect between ear biometric morphological variation and transformer architecture positional sensitivity, presenting a possible avenue for authentication schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。