arXiv:2605.16550cs.CVcs.LG2026-05中稿 · ICIP 2026

用注意力机制提升视频虹膜识别准确率,适合监控场景。

Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition

论文配图:Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
图 1 · 摘自论文原文
  • 引入注意力感知的Transformer聚合帧特征
  • 在COX Face数据集上达99.8%真阳性率
  • 适合低质量监控视频中的身份识别

视频虹膜识别是基于人眼周围区域识别个体身份的任务。该区域是人脸最具区分性的部位之一,特别适用于常规生物特征(如人脸或虹膜)因采集条件受限而失效的监控场景。本文提出一种面向监控环境的视频虹膜识别注意力感知方法。框架包含特征嵌入和特征聚合两个模块:特征嵌入采用深度卷积神经网络将虹膜数据映射为特征向量;特征聚合采用仅编码器的Transformer,自适应地将帧级特征聚合为单一视频表示及静态参考图像的特征向量。在公开的COX Face数据集上的实验表明,该方法显著优于简单聚合方案,在最佳情况下达到TPR@1e−1为99.8%,Rank-5为96.6%。

原文摘要 · Abstract (English)

Video periocular recognition is the task of recognizing an individual's identity based on the region around an individual's eyes. The periocular area is one of the most discriminative regions of the human face, making it suitable for recognition tasks. Its use as a biometric modality has emerged as an alternative, especially in surveillance scenarios where conventional biometric traits such as face or iris recognition become unfeasible due to unconstrained acquisition conditions. This paper proposes an attention-aware approach for video-based periocular recognition in surveillance environments. The framework consists of two main modules: feature embedding and aggregation. The feature embedding module is a deep convolutional neural network that maps periocular data to feature vectors. The aggregation module is an encoder-only transformer that adaptively learns to aggregate frame-level features into a single video representation and a feature vector for the still reference image. Experiments on the publicly available COX Face dataset show the robustness of the proposed method, consistently outperforming naive aggregation schemes. In the best scenario, the approach achieves $99.8\%$ of TPR@$1e^{-1}$ and $96.6\%$ of Rank-5.

视频识别注意力机制生物特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。