arXiv:2604.10541cs.CV2026-04中稿 · IEEE Transactions …

跨数据集双向学习面部动作单元与表情,提升细粒度与粗粒度识别效果。

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

论文配图:Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
图 1 · 摘自论文原文
  • 通过结构化语义映射框架,实现动作单元与表情的双向知识迁移。
  • 在多个基准上优于单任务和多任务基线,表情到动作单元的迁移有效提升性能。
  • 适合需要联合建模面部微表情与整体情绪的研究者使用。

面部动作单元(AU)检测与面部表情(FE)识别可视为情感面部行为任务,分别代表细微肌肉激活与整体情感状态。尽管两者存在内在语义关联,现有研究多集中于从AU到FE的知识迁移,双向学习仍不充分。实际中,这种挑战因异构数据条件而加剧:AU与FE数据集在标注范式(帧级与片段级)、标签粒度、数据可用性与多样性方面存在差异,阻碍了有效联合学习。为此,我们提出结构化语义映射(SSM)框架,在不同数据域与异构监督下实现双向AU-FE学习。SSM包含三个关键组件:(1) 共享视觉主干,从动态的AU与FE视频中学习统一的面部表示;(2) 通过文本语义原型(TSP)模块,利用固定文本描述构建可学习上下文提示的结构化语义原型,实现共享语义空间中的监督与跨任务对齐;(3) 动态先验映射(DPM)模块,融合FACS先验知识,并在文本语义空间中学习数据自适应的双向关联矩阵,实现显式知识传递。在主流的AU检测与FE识别基准上的大量实验表明,SSM持续优于单任务与多任务基线,性能媲美专用方法。FE到AU的结果进一步表明,整体表情语义可在异构数据集间为细粒度AU学习提供有效监督。

原文摘要 · Abstract (English)

Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations and coarse-grained holistic affective states, respectively. Despite their inherent semantic correlation, existing studies predominantly focus on knowledge transfer from AUs to FEs, while bidirectional learning remains insufficiently explored. In practice, this challenge is further compounded by heterogeneous data conditions, where AU and FE datasets differ in annotation paradigms (frame-level vs.\ clip-level), label granularity, and data availability and diversity, hindering effective joint learning. To address these issues, we propose a Structured Semantic Mapping (SSM) framework for bidirectional AU--FE learning under different data domains and heterogeneous supervision. SSM consists of three key components: (1) a shared visual backbone that learns unified facial representations from dynamic AU and FE videos; (2) semantic mediation via a Textual Semantic Prototype (TSP) module, which constructs structured semantic prototypes from fixed textual descriptions with learnable context prompts for supervision and cross-task alignment in a shared semantic space; and (3) a Dynamic Prior Mapping (DPM) module that incorporates FACS-derived prior knowledge and learns data-adaptive bidirectional association matrices in the textual semantic space for explicit knowledge transfer. Extensive experiments on popular AU detection and FE recognition benchmarks show that SSM consistently outperforms its single-task and multi-task baselines and achieves competitive performance against task-specific methods. The FE-to-AU results further show that holistic expression semantics provides useful supervision for fine-grained AU learning across heterogeneous datasets.

面部识别双向学习语义映射多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。