用跨模态融合与自监督学习,提升事件相机下人脸关键点定位精度。
Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- 通过跨模态融合关注机制,结合RGB信息增强事件数据特征提取
- 在真实事件数据集E-SIE上达到优于现有方法的准确率
- 适合需要低光、高速场景下人脸检测的研究者
事件相机在低光照和快速运动条件下具有高时间分辨率和光照鲁棒性,适用于人脸关键点对齐。然而,现有基于RGB的方法在事件数据上表现不佳,而仅使用事件数据训练常因空间信息有限导致性能不足。此外,缺乏全面标注的事件数据集也制约了该领域发展。为此,我们提出一种基于跨模态融合注意力(CMFA)与自监督多事件表征学习(SSMER)的新框架。CMFA通过引入对应RGB数据,引导模型从事件输入中提取鲁棒的人脸特征;SSMER则利用无标签事件数据实现有效特征学习,缓解空间信息缺失问题。在真实事件数据集E-SIE及公开WFLW-V的合成事件版本上的大量实验表明,本方法在多个评估指标上持续超越当前最优方法。
原文摘要 · Abstract (English)
Event cameras offer unique advantages for facial keypoint alignment under challenging conditions, such as low light and rapid motion, due to their high temporal resolution and robustness to varying illumination. However, existing RGB facial keypoint alignment methods do not perform well on event data, and training solely on event data often leads to suboptimal performance because of its limited spatial information. Moreover, the lack of comprehensive labeled event datasets further hinders progress in this area. To address these issues, we propose a novel framework based on cross-modal fusion attention (CMFA) and self-supervised multi-event representation learning (SSMER) for event-based facial keypoint alignment. Our framework employs CMFA to integrate corresponding RGB data, guiding the model to extract robust facial features from event input images. In parallel, SSMER enables effective feature learning from unlabeled event data, overcoming spatial limitations. Extensive experiments on our real-event E-SIE dataset and a synthetic-event version of the public WFLW-V benchmark show that our approach consistently surpasses state-of-the-art methods across multiple evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。