arXiv:2512.01145cs.CV2025-12

仅用时间标记实现微表情强度连续估计,无需逐帧标注。

Weakly Supervised Continuous Micro-Expression Intensity Estimation Using Temporal Deep Neural Network

  • 用三角先验将稀疏时间点转为密集伪强度轨迹,结合轻量时序模型预测每帧强度。
  • 在SAMM数据集上达到0.9014的斯皮尔曼相关系数,CASME II上最高达0.9116。
  • 适合关注情绪动态演化、缺乏逐帧标注资源的研究者使用。

微表情是短暂且不由自主的面部动作,反映真实情绪状态。现有研究多聚焦于离散类别分类,较少关注强度随时间的连续变化。该进展受限于缺乏逐帧强度标签,使全监督回归不可行。本文提出一种统一框架,仅需弱时间标签(起始、峰值、结束)即可进行连续微表情强度估计。通过简单三角先验将稀疏时间地标转换为密集伪强度轨迹,并采用基于ResNet18编码器与双向GRU的轻量时序回归模型,直接从图像序列预测帧级强度。方法无需逐帧标注,且通过统一预处理与时间对齐流程适配多个数据集。在SAMM和CASME II上的实验表明,模型与伪强度轨迹具有强时间一致性:在SAMM上,斯皮尔曼相关达0.9014,肯德尔相关达0.7999,优于逐帧基线;在CASME II上,不使用峰值排序项时分别达到0.9116和0.8168。消融实验证明,时序建模与结构化伪标签对捕捉微表情上升-峰值-下降动态至关重要。据我们所知,这是首个仅依赖稀疏时间标注的连续微表情强度估计统一方法。

原文摘要 · Abstract (English)

Micro-facial expressions are brief and involuntary facial movements that reflect genuine emotional states. While most prior work focuses on classifying discrete micro-expression categories, far fewer studies address the continuous evolution of intensity over time. Progress in this direction is limited by the lack of frame-level intensity labels, which makes fully supervised regression impractical. We propose a unified framework for continuous micro-expression intensity estimation using only weak temporal labels (onset, apex, offset). A simple triangular prior converts sparse temporal landmarks into dense pseudo-intensity trajectories, and a lightweight temporal regression model that combines a ResNet18 encoder with a bidirectional GRU predicts frame-wise intensity directly from image sequences. The method requires no frame-level annotation effort and is applied consistently across datasets through a single preprocessing and temporal alignment pipeline. Experiments on SAMM and CASME II show strong temporal agreement with the pseudo-intensity trajectories. On SAMM, the model reaches a Spearman correlation of 0.9014 and a Kendall correlation of 0.7999, outperforming a frame-wise baseline. On CASME II, it achieves up to 0.9116 and 0.8168, respectively, when trained without the apex-ranking term. Ablation studies confirm that temporal modeling and structured pseudo labels are central to capturing the rise-apex-fall dynamics of micro-facial movements. To our knowledge, this is the first unified approach for continuous micro-expression intensity estimation using only sparse temporal annotations.

微表情时序建模弱监督情感计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。