通过细粒度对齐文本与视觉线索,提升动态表情识别准确率。
From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition
- 分步优化文本描述,精准捕捉情绪语义
- 用运动差异加权突出相关面部动作,抑制干扰
- 在模糊或不均衡数据上表现更优,适合真实场景
动态面部表情识别(DFER)旨在从随时间变化的面部动作中识别情绪,是情感计算的关键。现有视觉-语言方法虽引入文本描述指导识别,但仍存在两大缺陷:未充分利用生成文本中的细微情绪线索,且缺乏有效机制过滤与情绪无关的面部动态。为此,我们提出GRACE框架,通过动态建模、语义文本精炼和词元级跨模态对齐,实现情绪显著时空特征的精准定位。该方法利用粗到细的情绪文本增强(CATE)模块生成情绪感知文本描述,并通过运动差值加权机制突出表情相关面部运动。这些优化后的语义与视觉信号采用熵正则化最优传输在词元级别对齐。在三个基准数据集上的实验表明,本方法显著提升识别性能,尤其在情绪模糊或类别不平衡等挑战性场景下,于平均召回率(UAR)和加权平均召回率(WAR)上均达到新SOTA水平。
原文摘要 · Abstract (English)
Dynamic Facial Expression Recognition (DFER) aims to identify human emotions from temporally evolving facial movements and plays a critical role in affective computing. While recent vision-language approaches have introduced semantic textual descriptions to guide expression recognition, existing methods still face two key limitations: they often underutilize the subtle emotional cues embedded in generated text, and they have yet to incorporate sufficiently effective mechanisms for filtering out facial dynamics that are irrelevant to emotional expression. To address these gaps, We propose GRACE, Granular Representation Alignment for Cross-modal Emotion recognition that integrates dynamic motion modeling, semantic text refinement, and token-level cross-modal alignment to facilitate the precise localization of emotionally salient spatiotemporal features. Our method constructs emotion-aware textual descriptions via a Coarse-to-fine Affective Text Enhancement (CATE) module and highlights expression-relevant facial motion through a motion-difference weighting mechanism. These refined semantic and visual signals are aligned at the token level using entropy-regularized optimal transport. Experiments on three benchmark datasets demonstrate that our method significantly improves recognition performance, particularly in challenging settings with ambiguous or imbalanced emotion classes, establishing new state-of-the-art (SOTA) results in terms of both UAR and WAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。