arXiv:2412.13168cs.CVcs.AI2024-12被引 2

隐式分离表情动态,提升野外场景识别准确率

Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild

  • 基于可学习的小波提升框架,隐式解耦情绪相关与无关的面部动态
  • 在多个野外数据集上超越现有监督方法,准确率更高且效率相当
  • 无需外部引导或显式操作,适合复杂真实场景的表情识别任务

野外动态面部表情识别(DFER)面临情绪相关表达被无关动作和全局上下文稀释的挑战。现有方法多依赖耦合时空表征,易引入非情绪相关特征偏差。本文提出隐式面部动态解耦框架(IFDD),通过扩展小波提升方案实现完全可学习的隐式解耦。该过程分两阶段:第一阶段为帧间静态-动态分割模块(ISSM),利用帧间相关性生成内容感知的分割索引,将特征分为具高全局相似性的组与具独特动态特性的组;第二阶段为基于提升的聚合-解耦模块(LADM),通过更新器聚合两组特征获得细粒度全局上下文特征,并由预测器从中解耦出情绪相关动态特征。大量实验表明,IFDD在多个野外数据集上优于现有监督方法,识别准确率更高,计算效率相当。代码已开源。

原文摘要 · Abstract (English)

In-the-wild dynamic facial expression recognition (DFER) encounters a significant challenge in recognizing emotion-related expressions, which are often temporally and spatially diluted by emotion-irrelevant expressions and global context. Most prior DFER methods directly utilize coupled spatiotemporal representations that may incorporate weakly relevant features with emotion-irrelevant context bias. Several DFER methods highlight dynamic information for DFER, but following explicit guidance that may be vulnerable to irrelevant motion. In this paper, we propose a novel Implicit Facial Dynamics Disentanglement framework (IFDD). Through expanding wavelet lifting scheme to fully learnable framework, IFDD disentangles emotion-related dynamic information from emotion-irrelevant global context in an implicit manner, i.e., without exploit operations and external guidance. The disentanglement process contains two stages. The first is Inter-frame Static-dynamic Splitting Module (ISSM) for rough disentanglement estimation, which explores inter-frame correlation to generate content-aware splitting indexes on-the-fly. We utilize these indexes to split frame features into two groups, one with greater global similarity, and the other with more unique dynamic features. The second stage is Lifting-based Aggregation-Disentanglement Module (LADM) for further refinement. LADM first aggregates two groups of features from ISSM to obtain fine-grained global context features by an updater, and then disentangles emotion-related facial dynamic features from the global context by a predictor. Extensive experiments on in-the-wild datasets have demonstrated that IFDD outperforms prior supervised DFER methods with higher recognition accuracy and comparable efficiency. Code is available at https://github.com/CyberPegasus/IFDD.

表情识别动态解耦小波提升野外场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。