用分层对齐与状态空间模型,提升真实场景下微表情检测精度。
Hierarchical Granularity Alignment and State Space Modeling for Robust Multimodal AU Detection in the Wild
- 分层对齐动态匹配全局表情与局部面部变化,应对极端姿态差异。
- 引入Vision-Mamba实现线性复杂度建模,捕捉超长时序依赖。
- 适合做真实环境下的多模态情感分析研究者参考。
真实场景下的面部动作单元(AU)检测面临严重的时空异质性、自由姿态和复杂的音视频依赖问题。现有方法多依赖容量受限的编码器和浅层融合机制,难以捕捉细粒度语义变化和超长时序上下文。为此,本文提出一种基于分层粒度对齐与状态空间模型的新型多模态框架。我们采用DINOv2和WavLM等基础模型提取鲁棒且高保真的视觉与音频表征,替代传统特征提取器。为应对极端面部变化,提出的分层粒度对齐模块可动态对齐全局面部语义与局部激活区域。此外,通过引入Vision-Mamba架构克服传统卷积网络感受野限制,实现O(N)线性复杂度的时间建模,有效捕捉超长动态而无性能下降。同时设计了一种非对称交叉注意力机制,深度同步副语言音频线索与细微视觉运动。在挑战性数据集Aff-Wild2上的大量实验表明,该方法显著优于现有基线,达到当前最优水平,并在第10届情感行为分析在野竞赛的AU检测赛道中获得第一名。
原文摘要 · Abstract (English)
Facial Action Unit (AU) detection in in-the-wild environments remains a formidable challenge due to severe spatial-temporal heterogeneity, unconstrained poses, and complex audio-visual dependencies. While recent multimodal approaches have made progress, they often rely on capacity-limited encoders and shallow fusion mechanisms that fail to capture fine-grained semantic shifts and ultra-long temporal contexts. To bridge this gap, we propose a novel multimodal framework driven by Hierarchical Granularity Alignment and State Space Models.Specifically, we leverage powerful foundation models, namely DINOv2 and WavLM, to extract robust and high-fidelity visual and audio representations, effectively replacing traditional feature extractors. To handle extreme facial variations, our Hierarchical Granularity Alignment module dynamically aligns global facial semantics with fine-grained local active patches. Furthermore, we overcome the receptive field limitations of conventional temporal convolutional networks by introducing a Vision-Mamba architecture. This approach enables temporal modeling with O(N) linear complexity, effectively capturing ultra-long-range dynamics without performance degradation. A novel asymmetric cross-attention mechanism is also introduced to deeply synchronize paralinguistic audio cues with subtle visual movements.Extensive experiments on the challenging Aff-Wild2 dataset demonstrate that our approach significantly outperforms existing baselines, achieving state-of-the-art performance. Notably, this framework secured top rankings in the AU Detection track of the 10th Affective Behavior Analysis in-the-wild Competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。