用事件相机实现鲁棒人脸识别,关键在建模人脸结构与动态特征。
EventFace: Event-Based Face Recognition via Structure-Driven Spatiotemporal Modeling
- 基于人脸刚性运动和几何结构,融合时空特征建模身份
- 在自建数据集上达到94.19%的识别率和5.35%的错误率
- 适合光照变化大或隐私敏感场景下的身份识别应用
事件相机因其对光照不敏感和隐私友好等优势,为人脸识别提供了新可能。由于事件流缺乏传统RGB系统依赖的稳定光度外观,本文主张通过建模由刚性面部运动和个体面部几何决定的结构驱动时空身份表示来实现事件相机下的人脸识别。由于缺乏专用数据集,我们构建了小规模事件人脸数据集EFace,其在刚性面部运动下采集。为有效利用有限事件数据,我们提出EventFace框架,整合空间结构与时间动态进行身份建模。具体地,采用低秩适配(LoRA)将预训练RGB人脸模型中的面部先验迁移至事件域,建立可靠的空基。在此基础上,引入运动提示编码器(MPE)显式编码时间特征,并设计时空调制器(STM)融合时空特征,增强身份相关事件模式的表示能力。大量实验表明,EventFace在所评估基线中表现最优,达到94.19%的排名1识别率和5.35%的等错误率(EER)。结果还显示,该方法在光照退化条件下具有更强鲁棒性,且学习到的表征可重构性降低。
原文摘要 · Abstract (English)
Event cameras offer a promising sensing modality for face recognition due to their inherent advantages in illumination robustness and privacy-friendliness. However, because event streams lack the stable photometric appearance relied upon by conventional RGB-based face recognition systems, we argue that event-based face recognition should model structure-driven spatiotemporal identity representations shaped by rigid facial motion and individual facial geometry. Since dedicated datasets for event-based face recognition remain lacking, we construct EFace, a small-scale event-based face dataset captured under rigid facial motion. To learn effectively from this limited event data, we further propose EventFace, a framework for event-based face recognition that integrates spatial structure and temporal dynamics for identity modeling. Specifically, we employ Low-Rank Adaptation (LoRA) to transfer structural facial priors from pretrained RGB face models to the event domain, thereby establishing a reliable spatial basis for identity modeling. Building on this foundation, we further introduce a Motion Prompt Encoder (MPE) to explicitly encode temporal features and a Spatiotemporal Modulator (STM) to fuse them with spatial features, thereby enhancing the representation of identity-relevant event patterns. Extensive experiments demonstrate that EventFace achieves the best performance among the evaluated baselines, with a Rank-1 identification rate of 94.19% and an equal error rate (EER) of 5.35%. Results further indicate that EventFace exhibits stronger robustness under degraded illumination than the competing methods. In addition, the learned representations exhibit reduced template reconstructability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。