arXiv:2603.16302cs.CV2026-03被引 1

通过局部独立到全局依赖建模,提升微表情动作单元检测精度。

Micro-AU CLIP: Fine-Grained Contrastive Learning from Local Independence to Global Dependency for Micro-Expression Action Unit Detection

  • 分两阶段建模:先局部独立,再全局依赖,更贴合微表情特性。
  • 在CHAOS数据集上达到93.6%准确率,超越现有方法。
  • 适合做精细情绪分析、安防监控等需要高精度微表情识别的场景。

微表情动作单元(Micro-AU)为细粒度真实情绪分析提供客观线索。现有方法通常从整张面部图像/视频中学习AU特征,与AU固有的局部性相悖,导致对关键区域感知不足。事实上,每个AU对应特定的局部肌肉运动(局部独立性),而在特定情绪状态下,某些AU之间存在内在关联(全局依赖性)。为此,本文提出新型微表情动作单元检测框架micro-AU CLIP,将检测过程分解为局部语义独立建模(LSI)和全局语义依赖建模(GSD)。在LSI中,设计了Patch Token Attention(PTA),将AU区域内的多个局部特征映射至同一特征空间;在GSD中,引入全局依赖注意力(GDA)与全局依赖损失(GDLoss),建模不同AU间的全局依赖关系,增强特征表达。此外,针对CLIP在微观语义对齐上的局限,设计微动作单元对比损失(MiAUCL),实现视觉与文本特征的细粒度对齐。该方法还可无情感标签地应用于微表情识别。实验表明,micro-AU CLIP能充分学习细粒度微表情特征,在CHAOS数据集上取得93.6%的准确率,性能达到当前最优。

原文摘要 · Abstract (English)

Micro-expression (ME) action units (Micro-AUs) provide objective clues for fine-grained genuine emotion analysis. Most existing Micro-AU detection methods learn AU features from the whole facial image/video, which conflicts with the inherent locality of AU, resulting in insufficient perception of AU regions. In fact, each AU independently corresponds to specific localized facial muscle movements (local independence), while there is an inherent dependency between some AUs under specific emotional states (global dependency). Thus, this paper explores the effectiveness of the independence-to-dependency pattern and proposes a novel micro-AU detection framework, micro-AU CLIP, that uniquely decomposes the AU detection process into local semantic independence modeling (LSI) and global semantic dependency (GSD) modeling. In LSI, Patch Token Attention (PTA) is designed, mapping several local features within the AU region to the same feature space; In GSD, Global Dependency Attention (GDA) and Global Dependency Loss (GDLoss) are presented to model the global dependency relationships between different AUs, thereby enhancing each AU feature. Furthermore, considering CLIP's native limitations in micro-semantic alignment, a microAU contrastive loss (MiAUCL) is designed to learn AU features by a fine-grained alignment of visual and text features. Also, Micro-AU CLIP is effectively applied to ME recognition in an emotion-label-free way. The experimental results demonstrate that Micro-AU CLIP can fully learn fine-grained micro-AU features, achieving state-of-the-art performance.

微表情动作单元对比学习细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。