用因果监督提升注意力机制,让情绪分析更准确且泛化更强。
Causal Supervision of Attention for Affective Behaviour Analysis

- 引入因果监督与交叉协方差正则化,让注意力聚焦真实情绪线索。
- 在多个任务上取得0.5123(VA)、0.3116(EX)、0.3974(AU)的指标表现。
- 适合关注情绪识别泛化能力与注意力可解释性的研究者。
第11届野外情感行为分析竞赛包含多任务学习挑战,要求构建统一框架完成愉悦度-唤醒度估计、表情识别和动作单元检测。难点在于学习跨受试者泛化的感情特征,同时对身份、光照、姿态及人口统计差异等伪相关因素保持鲁棒。为将预训练主干网络提取的特征聚合为紧凑表示以进行预测,注意力机制会加权最信息丰富的面部区域。然而,这些注意力权重仍可能捕捉数据集特有相关性而非真实情感信号。为此,我们提出一种结合因果监督与注意力组件交叉协方差正则化的注意力池化框架,促进主体不变的注意力分布和非冗余表示,从而提升泛化性能。在官方验证集上,该方法实现愉悦度-唤醒度估计的CCC_{VA}=0.5123,表情识别的F_{EX}=0.3116,动作单元检测的F_{AU}=0.3974,总分P=1.2214。
原文摘要 · Abstract (English)
The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves $CCC_{VA}=0.5123$ for VA estimation on the official validation set, together with $F_{EX}=0.3116$ and $F_{AU}=0.3974$ for expression recognition and action unit detection, respectively, resulting in an overall $P$ score (the sum of the individual task metrics) of $1.2214$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。