arXiv:2604.22842cs.CVeess.IV2026-04中稿 · FG2026

利用视觉变压器中间层特征提升人脸图像质量评估精度

EX-FIQA: Leveraging Intermediate Early eXit Representations from Vision Transformers for Face Image Quality Assessment

论文配图:EX-FIQA: Leveraging Intermediate Early eXit Representations from Vision Transformers for Face Image Quality Assessment
图 1 · 摘自论文原文
  • 通过早期退出机制捕获ViT各层中间特征,融合多层质量预测
  • 在8个数据集上实现比单层输出更高性能,计算量减少30%以上
  • 无需修改模型或重新训练,适合资源受限的生物识别系统

人脸图像质量评估对可靠的人脸识别系统至关重要,但现有的基于视觉变换器(ViT)的方法仅依赖最终层表示,忽略了中间网络深度中捕捉的质量相关信息。本文首次全面研究了ViT内部中间表示如何通过早期退出机制和分数融合策略贡献于人脸质量评估。我们系统分析了ViT-FIQA架构中全部12个Transformer块,证明不同深度捕捉到具有差异性和互补性的质量相关特征,表现为各层注意力模式和性能特征的差异。我们提出一种无需架构修改或额外训练的分数融合框架,结合多个Transformer块的质量预测。早期退出分析揭示了性能与效率的最佳权衡,实现显著计算节省的同时保持竞争力。在四个全参考(FR)模型下,跨八个基准数据集的广泛评估表明,我们的融合策略优于单一出口方法。所提出的质量融合方法采用深度加权平均,逐步提高深层块的重要性,有效利用了ViT中特征学习的层次结构。本工作挑战了‘仅深层特征重要’的常规认知,揭示中间表示包含评估质量的宝贵信息。该框架为实际生物识别系统提供了实用优势,支持根据资源约束自适应计算,同时维持良好的质量评估能力。

原文摘要 · Abstract (English)

Face Image Quality Assessment is crucial for reliable face recognition systems, yet existing Vision Transformer-based approaches rely exclusively on final-layer representations, ignoring quality-relevant information captured at intermediate network depths. This paper presents the first comprehensive investigation of how intermediate representations within ViTs contribute to face quality assessment through early exit mechanisms and score fusion strategies. We systematically analyze all twelve transformer blocks of ViT-FIQA architectures, demonstrating that different depths capture distinct and complementary quality-relevant information, as evidenced by varying attention patterns and performance characteristics across network layers. We propose a score fusion framework that combines quality predictions from multiple transformer blocks without architectural modifications or additional training. Our early exit analysis reveals optimal performance-efficiency trade-offs, enabling significant computational savings while maintaining competitive performance. Through extensive evaluation across eight benchmark datasets using four FR models, we demonstrate that our fusion strategy improves upon single-exit approaches. Our proposed quality fusion approach employs depth-weighted averaging that assigns progressively higher importance to deeper transformer blocks, achieving the best quality assessment performance by effectively leveraging the hierarchical nature of feature learning in ViTs. Our work challenges the conventional wisdom that only deep features matter for face analysis, revealing that intermediate representations contain valuable information for quality assessment. The proposed framework offers practical benefits for real-world biometric systems by enabling adaptive computation based on resource constraints while maintaining competitive quality assessment capabilities.

人脸质量评估视觉变换器早期退出特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。