用语言描述提升面部动作单元检测,更准更可解释。
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
- 用大模型生成多样文本描述,引导视觉特征学习
- 设计双注意力机制,捕捉局部与全局的跨模态关联
- 在3个数据集上超越现有方法,结果可解释性强
面部动作单元(AU)检测旨在识别由面部动作编码系统(FACS)定义的细微肌肉激活。其主要挑战在于有限标注数据下学习判别性与泛化性强的AU表征。为此,我们提出一种分层视觉-语言交互方法(HiVA),利用文本形式的AU描述作为语义先验,指导并增强AU检测。HiVA通过大语言模型生成丰富多样的上下文描述,强化语言表征学习;引入面向AU的动态图模块,捕获细粒度与整体的视觉-语言关联;并在分层跨模态注意力架构中集成两种机制:解耦双交叉注意力(DDCA)实现细粒度、特定于AU的视觉-文本交互,上下文双交叉注意力(CDCA)建模全局跨AU依赖。该协同跨模态学习范式使HiVA结合多粒度视觉特征与精细化语言信息,实现鲁棒且语义丰富的AU检测。大量实验表明,HiVA持续优于当前最优方法;定性分析显示其生成语义合理的激活模式,验证了在全面面部行为分析中学习强而可解释的跨模态对应关系的有效性。
原文摘要 · Abstract (English)
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Interaction for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal attention architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross-modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language-based AU details, culminating in robust and semantically enriched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。