arXiv:2511.10958cs.CVcs.AI2025-11被引 1

用文本指导弱监督框架,提升动态表情识别的准确性与可解释性

Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition

  • 引入视觉-语言模型提供细粒度情感描述语义引导
  • 通过视觉提示对齐文本标签与图像特征,实现帧级相关性判断
  • 多粒度时序网络捕捉短期动作与长期情绪流,增强时间敏感性

动态面部表情识别(DFER)旨在通过视频序列中面部运动的时间变化来识别情绪状态。其核心挑战在于‘多对一’标注问题:一段包含多个帧的视频仅被赋予单一情绪标签。常见解决策略是将DFER建模为多实例学习(MIL)问题,但此类方法受情绪表达视觉多样性与时序复杂性影响,性能受限。为此,本文提出TG-DFER——一种文本引导的弱监督框架,通过引入视觉-语言预训练(VLP)模型,利用细粒度文本描述提供语义指导;同时设计视觉提示机制,将丰富的情感标签与视觉实例特征对齐,支持细粒度推理与帧级相关性估计;此外,构建多粒度时序网络,联合捕捉短期面部动态与长期情绪流,实现时间上的一致性情感理解。大量实验表明,该方法在弱监督下显著提升了泛化能力、可解释性与时间敏感性。

原文摘要 · Abstract (English)

Dynamic facial expression recognition (DFER) aims to identify emotional states by modeling the temporal changes in facial movements across video sequences. A key challenge in DFER is the many-to-one labeling problem, where a video composed of numerous frames is assigned a single emotion label. A common strategy to mitigate this issue is to formulate DFER as a Multiple Instance Learning (MIL) problem. However, MIL-based approaches inherently suffer from the visual diversity of emotional expressions and the complexity of temporal dynamics. To address this challenge, we propose TG-DFER, a text-guided weakly supervised framework that enhances MIL-based DFER by incorporating semantic guidance and coherent temporal modeling. We incorporate a vision-language pre-trained (VLP) model is integrated to provide semantic guidance through fine-grained textual descriptions of emotional context. Furthermore, we introduce visual prompts, which align enriched textual emotion labels with visual instance features, enabling fine-grained reasoning and frame-level relevance estimation. In addition, a multi-grained temporal network is designed to jointly capture short-term facial dynamics and long-range emotional flow, ensuring coherent affective understanding across time. Extensive results demonstrate that TG-DFER achieves improved generalization, interpretability, and temporal sensitivity under weak supervision.

表情识别弱监督视觉语言模型时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。