无监督跨域表情识别新框架,提升模型在低资源场景下的泛化能力。
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
- 用视觉语言模型对齐情绪表征,引导扩散模型增强感知
- 通过反事实生成伪标签,解决目标域情感分布偏移问题
- 适用于真实图像到贴纸等低资源场景的表情识别任务
视觉表情识别(VER)旨在基于视觉线索推断个体情绪状态,但现有方法多局限于单一领域(如真实图像或贴纸),限制了模型的跨域泛化能力。为此,我们提出无监督跨域视觉表情识别(UCDVER)任务,旨在将源域(如真实图像)的情绪识别能力迁移至低资源目标域(如贴纸),且无需目标域标注。与传统无监督域适应不同,UCDVER面临两大挑战:情绪表达差异大、情感分布显著偏移。为此,我们提出知识对齐的反事实增强扩散感知框架(KCDP)。该框架利用视觉语言模型(VLM)将情绪表征映射到共享知识空间,并指导扩散模型提升情绪感知能力;同时引入反事实增强的语言-图像情绪对齐(CLIEA)方法,为目标域生成高质量伪标签。大量实验表明,本模型在可感知性和泛化能力上均优于当前最优模型,在多个指标上相较SOTA模型TGCA-PVT提升12%。项目主页:https://yinwen2019.github.io/ucdver。
原文摘要 · Abstract (English)
Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recognition (UCDVER) task, which aims to generalize visual emotion recognition from the source domain (e.g., realistic images) to the low-resource target domain (e.g., stickers) in an unsupervised manner. Compared to the conventional unsupervised domain adaptation problems, UCDVER presents two key challenges: a significant emotional expression variability and an affective distribution shift. To mitigate these issues, we propose the Knowledge-aligned Counterfactual-enhancement Diffusion Perception (KCDP) framework. Specifically, KCDP leverages a VLM to align emotional representations in a shared knowledge space and guides diffusion models for improved visual affective perception. Furthermore, a Counterfactual-Enhanced Language-image Emotional Alignment (CLIEA) method generates high-quality pseudo-labels for the target domain. Extensive experiments demonstrate that our model surpasses SOTA models in both perceptibility and generalization, e.g., gaining 12% improvements over the SOTA VER model TGCA-PVT. The project page is at https://yinwen2019.github.io/ucdver.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。