arXiv:2410.21958cs.CV2024-10被引 2

用事件相机和时空变压器提升微表情识别准确率

Spatio-temporal Transformers for Action Unit Classification with Event Cameras

  • 提出基于SPT与LSA的时空视觉变换器模型,捕捉事件流中的时空特征
  • 在自建的多模态数据集FACEMORPHIC上实现91.3%的分类准确率,优于基线方法
  • 通过跨模态监督实现无需人工标注的神经形态面部分析,适合高帧率应用

面部分析常用于推断情绪、姿态、形状和关键点。传统RGB相机在精细任务中受限于延迟,难以记录和检测携带丰富信息的微运动,影响真实情绪推断。事件相机因高帧率潜力日益受到关注。本文提出一种新型时空视觉变压器模型,采用移位块标记化(SPT)和局部自注意力(LSA),以提升从事件流中进行动作单元(Action Unit)分类的准确性。针对文献中缺乏标注事件数据的问题,我们构建了FACEMORPHIC数据集,包含同步的RGB视频与事件流,按视频级别标注面部动作单元,涵盖3D形状估计到唇读等多种应用场景。通过时间同步,利用跨模态监督将面部形状映射至三维空间,实现无需手动标注的神经形态面部分析。所提模型有效捕捉空间与时间信息,在识别细微面部微表情方面表现优异。

原文摘要 · Abstract (English)

Face analysis has been studied from different angles to infer emotion, poses, shapes, and landmarks. Traditionally RGB cameras are used, yet for fine-grained tasks standard sensors might not be up to the task due to their latency, making it impossible to record and detect micro-movements that carry a highly informative signal, which is necessary for inferring the true emotions of a subject. Event cameras have been increasingly gaining interest as a possible solution to this and similar high-frame rate tasks. We propose a novel spatiotemporal Vision Transformer model that uses Shifted Patch Tokenization (SPT) and Locality Self-Attention (LSA) to enhance the accuracy of Action Unit classification from event streams. We also address the lack of labeled event data in the literature, which can be considered one of the main causes of an existing gap between the maturity of RGB and neuromorphic vision models. Gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. To this end, we present FACEMORPHIC, a temporally synchronized multimodal face dataset composed of RGB videos and event streams. The dataset is annotated at a video level with facial Action Units and contains streams collected with various possible applications, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space. Our proposed model outperforms baseline methods by effectively capturing spatial and temporal information, crucial for recognizing subtle facial micro-expressions.

事件相机微表情视觉变换器多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。