用同步视频与事件数据,实现无需人工标注的神经形态面部分析。
Neuromorphic Facial Analysis with Cross-Modal Supervision
- 构建同步的RGB与事件流面部数据集,支持多任务应用。
- 通过时间对齐,实现跨模态监督,避免手动标注。
- 适合研究神经形态视觉与情感计算的学者。
传统基于RGB帧的方法能从不同角度精细分析人脸情绪、姿态、形状和关键点,但在捕捉细微动作时受限于相机延迟,难以检测携带重要情感信息的微运动。事件相机因其低延迟特性逐渐受到关注,但其数据表示方式与RGB存在显著差异,导致成熟RGB处理技术难以直接迁移。此外,事件域缺乏标注数据,且无法从网络爬取,标注需考虑事件聚合率及静态区域在特定帧中不可见的问题。本文提出FACEMORPHIC,一个包含同步RGB视频与事件流的多模态面部数据集,视频级标注了面部动作单元(Action Units),并支持3D形状估计、唇读等多样化应用。通过时间同步,可实现跨模态监督,将人脸形状映射至3D空间,有效缓解领域差异,无需人工标注即可进行神经形态面部分析。
原文摘要 · Abstract (English)
Traditional approaches for analyzing RGB frames are capable of providing a fine-grained understanding of a face from different angles by inferring emotions, poses, shapes, landmarks. However, when it comes to subtle movements standard RGB cameras might fall behind due to their latency, making it hard to detect micro-movements that carry highly informative cues to infer the true emotions of a subject. To address this issue, the usage of event cameras to analyze faces is gaining increasing interest. Nonetheless, all the expertise matured for RGB processing is not directly transferrable to neuromorphic data due to a strong domain shift and intrinsic differences in how data is represented. The lack of labeled data can be considered one of the main causes of this gap, yet gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. In this paper, we first present FACEMORPHIC, a multimodal temporally synchronized face dataset comprising both RGB videos and event streams. The data is labeled at a video level with facial Action Units and also contains streams collected with a variety of applications in mind, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。