arXiv:2502.21004cs.CV2025-02被引 1

用自适应时间软掩码提升动态表情识别效率与精度

Soften the Mask: Adaptive Temporal Soft Mask for Efficient Dynamic Facial Expression Recognition

  • 引入自适应时间软掩码,动态筛选关键表情帧
  • 计算成本显著降低,性能仍达领先水平
  • 适合需要高效实时表情识别的应用场景

动态面部表情识别(DFER)通过非语言交流理解心理意图。现有方法难以处理背景噪声和冗余语义等无关信息,影响效率与效果。本文提出新型监督式时序软掩码自编码网络AdaTosk,融合并行监督分类分支与自监督重建分支。重建分支采用随机二值硬掩码生成多样训练样本,促进可见标记中有效特征表示;分类分支则使用自适应时序软掩码,根据时间重要性灵活掩码可见标记。其两个关键组件——类无关与类语义软掩码——分别增强关键表情时刻并减少时间上的语义冗余。在多个主流基准上的大量实验表明,与当前最优方法相比,AdaTosk 显著降低计算成本,同时保持竞争力的性能。

原文摘要 · Abstract (English)

Dynamic Facial Expression Recognition (DFER) facilitates the understanding of psychological intentions through non-verbal communication. Existing methods struggle to manage irrelevant information, such as background noise and redundant semantics, which impacts both efficiency and effectiveness. In this work, we propose a novel supervised temporal soft masked autoencoder network for DFER, namely AdaTosk, which integrates a parallel supervised classification branch with the self-supervised reconstruction branch. The self-supervised reconstruction branch applies random binary hard mask to generate diverse training samples, encouraging meaningful feature representations in visible tokens. Meanwhile the classification branch employs an adaptive temporal soft mask to flexibly mask visible tokens based on their temporal significance. Its two key components, respectively of, class-agnostic and class-semantic soft masks, serve to enhance critical expression moments and reduce semantic redundancy over time. Extensive experiments conducted on widely-used benchmarks demonstrate that our AdaTosk remarkably reduces computational costs compared with current state-of-the-art methods while still maintaining competitive performance.

表情识别自编码器高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。