用事件相机捕捉微表情,比传统摄像头更准更快。
Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis
- 同步采集RGB与事件数据,构建多模态微表情数据集。
- 事件数据识别微表情准确率达51.23%,远超RGB的23.12%。
- 适合做智能交互、驾驶监控等需要快速反应的场景。
微表情分析在人机交互和驾驶员监测系统中有重要应用。传统RGB相机因时间分辨率有限且易受运动模糊影响,难以准确捕捉细微快速的面部动作。事件相机具备微秒级精度、高动态范围和低延迟优势。然而,公开的含动作单元(Action Units)的事件数据集仍很稀缺。本文提出一个初步的多分辨率、多模态微表情数据集,通过同步的RGB与事件相机在不同光照条件下采集。评估了两个基准任务:基于脉冲神经网络(Spiking Neural Networks)的动作单元分类,事件输入达51.23%准确率,远高于RGB的23.12%;使用条件变分自编码器进行帧重建,高分辨率事件输入下达到SSIM=0.8513、PSNR=26.89 dB。结果表明,事件数据可用于微表情识别与帧重建。
原文摘要 · Abstract (English)
Micro-expression analysis has applications in domains such as Human-Robot Interaction and Driver Monitoring Systems. Accurately capturing subtle and fast facial movements remains difficult when relying solely on RGB cameras, due to limitations in temporal resolution and sensitivity to motion blur. Event cameras offer an alternative, with microsecond-level precision, high dynamic range, and low latency. However, public datasets featuring event-based recordings of Action Units are still scarce. In this work, we introduce a novel, preliminary multi-resolution and multi-modal micro-expression dataset recorded with synchronized RGB and event cameras under variable lighting conditions. Two baseline tasks are evaluated to explore the spatial-temporal dynamics of micro-expressions: Action Unit classification using Spiking Neural Networks (51.23\% accuracy with events vs. 23.12\% with RGB), and frame reconstruction using Conditional Variational Autoencoders, achieving SSIM = 0.8513 and PSNR = 26.89 dB with high-resolution event input. These promising results show that event-based data can be used for micro-expression recognition and frame reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。