arXiv:2503.20771cs.CV2025-03被引 12

用中性视频无监督提升表情识别模型,解决目标数据缺失难题。

Disentangled Source-Free Personalization for Facial Expression Recognition with Neutral Target Data

  • 从目标用户的中性视频生成缺失的非中性表情数据
  • 通过解耦身份与表情特征,提升模型在新用户上的准确率
  • 适合医疗场景下隐私敏感的表情识别应用

面部表情识别(FER)在人机交互和健康诊断中至关重要,但受个体间表达差异影响。传统无源域适应(SFDA)需完整目标数据集,而医疗场景常难以获取全类别数据。本文提出一种基于中性控制视频的无监督域适应方法(DSFDA),利用仅含中性表情的短视频,端到端生成缺失的非中性表情数据,并通过解耦身份与表情特征实现模型自适应。同时引入自监督重建策略,保持身份与源表情一致性,显著提升模型在目标个体上的表现。实验表明,该方法在仅使用中性视频的情况下,相比基线模型在Expr-2020和MMI数据集上分别提升6.8%和4.3%的准确率。

原文摘要 · Abstract (English)

Facial Expression Recognition (FER) from videos is a crucial task in various application areas, such as human-computer interaction and health diagnosis and monitoring (e.g., assessing pain and depression). Beyond the challenges of recognizing subtle emotional or health states, the effectiveness of deep FER models is often hindered by the considerable inter-subject variability in expressions. Source-free (unsupervised) domain adaptation (SFDA) methods may be employed to adapt a pre-trained source model using only unlabeled target domain data, thereby avoiding data privacy, storage, and transmission issues. Typically, SFDA methods adapt to a target domain dataset corresponding to an entire population and assume it includes data from all recognition classes. However, collecting such comprehensive target data can be difficult or even impossible for FER in healthcare applications. In many real-world scenarios, it may be feasible to collect a short neutral control video (which displays only neutral expressions) from target subjects before deployment. These videos can be used to adapt a model to better handle the variability of expressions among subjects. This paper introduces the Disentangled SFDA (DSFDA) method to address the challenge posed by adapting models with missing target expression data. DSFDA leverages data from a neutral target control video for end-to-end generation and adaptation of target data with missing non-neutral data. Our method learns to disentangle features related to expressions and identity while generating the missing non-neutral expression data for the target subject, thereby enhancing model accuracy. Additionally, our self-supervision strategy improves model adaptation by reconstructing target images that maintain the same identity and source expression.

表情识别无监督学习医疗应用数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。