arXiv:2603.27999cs.CV2026-03被引 3

用面部动作单元精准建模微表情,实现无需微调的个性化情感识别。

CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition

  • 以面部动作单元作为结构化文本提示,融入CLIP模型增强细粒度表达理解。
  • 在三个视频数据集上优于现有CLIP与测试时自适应方法,准确率提升显著。
  • 适合需要快速适配新用户微表情特征的实时情感分析场景。

情感识别中的个性化对捕捉细微且个体差异的表情模式至关重要。尽管如CLIP等视觉语言模型在联合图文表征方面展现潜力,但现有方法或依赖对比预训练,或使用大模型生成描述性提示,存在噪声多、计算开销大、难以捕捉细粒度表情等问题。本文提出利用面部动作单元(AUs)作为结构化文本提示,在CLIP中建模细微面部表情。AUs编码表情背后的肌肉活动,提供局部且可解释的语义线索,提升表情识别鲁棒性。我们提出CLIP-AU,一种轻量级的时空学习方法,通过对齐AU提示与面部动态,学习通用的、不受个体影响的表示,实现无需微调的细粒度人脸识别。然而,该方法未考虑个体间微妙表达的差异。为此,我们进一步提出CLIP-AUTT,一种基于视频的测试时个性化方法,动态调整来自未知受试者的视频对应的AU提示。结合熵引导的时间窗口选择与提示调优,实现个体特异性适应同时保持时间一致性。在BioVid、StressID和BAH三个挑战性视频数据集上的实验表明,CLIP-AU与CLIP-AUTT均超越当前最优的CLIP基线与测试时自适应方法。

原文摘要 · Abstract (English)

Personalization in emotion recognition (ER) is essential for accurate interpretation of subtle and subject-specific expressive patterns. Recent advances in vision-language models (VLMs), such as CLIP, demonstrate strong potential for leveraging joint image-text representations in ER. However, existing CLIP-based methods either rely on CLIP's contrastive pretraining or use LLMs to generate descriptive text prompts, which can be noisy, computationally expensive, and often fail to capture fine-grained expressions, leading to degraded performance. In this work, Action Units (AUs) are leveraged as structured textual prompts within CLIP to model fine-grained facial expressions. AUs encode the subtle muscle activations underlying expressions, providing localized and interpretable semantic cues for more robust facial expression recognition (FER). We introduce CLIP-AU, a lightweight AU-guided temporal learning method that integrates interpretable AU semantics into CLIP. It learns generic, subject-agnostic representations by aligning AU prompts with facial dynamics, enabling fine-grained FER without CLIP fine-tuning or LLM-generated text supervision. Although CLIP-AU models fine-grained AU semantics, it does not adapt to subject-specific variability in subtle expressions. To address this limitation, we propose CLIP-AUTT, a video-based test-time personalization method that dynamically adapts AU prompts to videos from unseen subjects. By combining entropy-guided temporal window selection with prompt tuning, CLIP-AUTT enables subject-specific adaptation while preserving temporal consistency. Our experiments on three challenging video-based datasets, BioVid, StressID, and BAH, indicate that CLIP-AU and CLIP-AUTT outperform state-of-the-art CLIP-based FER and TTA methods.

情感识别面部动作单元测试时自适应视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。