arXiv:2511.16952cs.CV2025-11中稿 · publication in IEE…

仅用一个时间点标注实现表情定位,提升精度与泛化能力。

Point-Supervised Facial Expression Spotting with Gaussian-Based Instance-Adaptive Intensity Modeling

  • 用高斯分布建模表情强度,生成软标签替代硬标签。
  • 在SAMM-LV等数据集上性能超越全监督方法。
  • 适合标注资源稀缺的表情分析场景。

自动面部表情定位旨在从非剪裁视频中识别表情实例,对表情分析至关重要。现有方法主要依赖全监督学习,需耗费大量时间和人力的时间边界标注。本文研究点标注式表情定位(P-FES),每实例仅需一个时间戳标注即可训练。提出双分支框架:首先,为缓解硬伪标签混淆中性与不同强度表情的问题,设计基于高斯的实例自适应强度建模(GIM)模块,通过检测伪顶点帧、估计持续时间并构建实例级高斯分布,实现软伪标签分配,提供更可靠的强度监督;该模块用于优化无类别表达强度分支。其次,设计有类别感知的顶点分类分支,仅基于伪顶点帧区分宏表情与微表情。推理时两分支独立工作:无类别强度分支生成表达提议,有类别分支负责宏/微表情分类。此外,引入强度感知对比损失,通过对比中性帧与不同强度表达帧,增强特征判别性并抑制中性噪声。在SAMM-LV、CAS(ME)²、CAS(ME)³数据集上的大量实验验证了所提框架的有效性。

原文摘要 · Abstract (English)

Automatic facial expression spotting, which aims to identify facial expression instances in untrimmed videos, is crucial for facial expression analysis. Existing methods primarily focus on fully-supervised learning and rely on costly, time-consuming temporal boundary annotations. In this paper, we investigate point-supervised facial expression spotting (P-FES), where only a single timestamp annotation per instance is required for training. We propose a unique two-branch framework for P-FES. First, to mitigate the limitation of hard pseudo-labeling, which often confuses neutral and expression frames with various intensities, we propose a Gaussian-based instance-adaptive intensity modeling (GIM) module to model instance-level expression intensity distribution for soft pseudo-labeling. By detecting the pseudo-apex frame around each point label, estimating the duration, and constructing an instance-level Gaussian distribution, GIM assigns soft pseudo-labels to expression frames for more reliable intensity supervision. The GIM module is incorporated into our framework to optimize the class-agnostic expression intensity branch. Second, we design a class-aware apex classification branch that distinguishes macro- and micro-expressions solely based on their pseudo-apex frames. During inference, the two branches work independently: the class-agnostic expression intensity branch generates expression proposals, while the class-aware apex-classification branch is responsible for macro- and micro-expression classification. Furthermore, we introduce an intensity-aware contrastive loss to enhance discriminative feature learning and suppress neutral noise by contrasting neutral frames with expression frames with various intensities. Extensive experiments on the SAMM-LV, CAS(ME)$^2$, and CAS(ME)$^3$ datasets demonstrate the effectiveness of our proposed framework.

表情识别弱监督高斯建模视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。