arXiv:2606.25542cs.CV2026-06

提出新模型融合面部微表情的运动与外观信息,提升识别准确率。

SAC$^2$-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition

论文配图:SAC$^2$-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition
图 1 · 摘自论文原文
  • 用语义锚点对齐不同模态的特征,保持解剖相关样本语义相近。
  • 通过互补共识融合机制,自适应增强不可靠局部响应,提升判别力。
  • 适用于小样本、细微表情识别任务,尤其适合跨模态融合场景。

微表情识别(MER)因面部动作细微、数据有限,且动作单元(AUs)与情绪类别关系模糊而具有挑战性。光学流和运动放大是常用表征,分别捕捉肌肉位移级运动与面部纹理及上下文变化。现有方法多将二者视为独立线索或简单融合,未显式建模其双重互补性:当两模态均有效时,联合提供更完整动态表征;当一模态退化时,另一模态仍可提供补偿性证据。为此,我们提出SAC²-Net,即语义锚定与互补共识融合网络,先通过语义锚点对齐异构视觉表示,再进行可靠性感知的互补融合。具体地,语义锚定软对齐(SASA)将激活的AUs转化为文本提示,利用分层的AU感知软标签对齐运动放大与光流特征,同时保持解剖相关样本的语义接近性。基于对齐特征,互补共识融合(CCF)交换互补的运动与外观线索,自适应增强不可靠局部响应,并通过共识精炼促进共享空间关注。该方法有效应对跨模态异质性与空间上可变的模态可靠性问题。

原文摘要 · Abstract (English)

Micro-expression recognition (MER) is challenging due to subtle facial movements, limited data, and the ambiguous relationship between Action Units (AUs) and emotion categories. Optical flow and motion magnification are two widely used representations for making subtle facial dynamics observable. However, many existing methods treat them as separate cues or fuse them without explicitly modeling their dual complementarity. Optical flow encodes displacement-level muscle motion, whereas motion magnification reveals appearance-level changes in facial texture and context. When both modalities are informative, their combination provides a more complete characterization of subtle facial dynamics; when one modality degrades, the other may still preserve discriminative evidence for compensation. This dual complementarity provides richer facial representations, but also introduces two key challenges for multimodal fusion: cross-modal heterogeneity and spatially varying modality reliability. To address these challenges, we propose SAC$^2$-Net, a Semantic Anchoring and Complementary-Consensus Network that first aligns heterogeneous visual representations with semantic anchors and then performs reliability-aware complementary fusion. Specifically, Semantic Anchoring Soft Alignment (SASA) converts activated AUs into textual prompts and uses hierarchical AU-aware soft labels to align motion-magnified and optical-flow representations while preserving semantic proximity among anatomically related samples. Based on the aligned representations, Complementary-Consensus Fusion (CCF) exchanges complementary motion and appearance cues, adaptively enhances unreliable local responses with trustworthy cross-modal evidence, and further encourages a shared spatial focus through consensus refinement.

微表情识别多模态融合语义对齐互补学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。