基于心理理论的跨模态预训练框架,提升场景情绪识别泛化能力
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
- 融合场景与人物的视觉结构信息,结合图文相似性增强情绪表征
- 在6个基准数据集上实现领先性能,Emo8数据集覆盖5类风格
- 适合研究跨场景情绪分析与多模态学习的学者使用
视觉情绪分析在计算机视觉与心理学领域具有重要价值。然而,现有方法因情绪感知模糊性和数据场景多样性,泛化能力受限。为此,我们提出UniEmoX,一种跨模态语义引导的大规模预训练框架。受心理学研究启发——情绪体验与个体环境互动密不可分,UniEmoX整合了以场景为中心和以人为中心的低级图像空间结构信息,旨在提取更细致、更具区分性的表情特征。通过利用配对与非配对图像-文本样本间的相似性,UniEmoX从CLIP模型中提炼丰富的语义知识,更有效地增强情绪嵌入表示。据我们所知,这是首个将心理理论与对比学习、掩码图像建模技术结合的大型预训练框架,用于跨多样化场景的情绪分析。此外,我们构建了名为Emo8的视觉情绪数据集,涵盖卡通、自然、写实、科幻及广告等多种风格,几乎覆盖所有常见情绪场景。在六个基准数据集上的全面实验验证了UniEmoX的有效性。源代码已公开于https://github.com/chincharles/u-emo。
原文摘要 · Abstract (English)
Visual emotion analysis holds significant research value in both computer vision and psychology. However, existing methods for visual emotion analysis suffer from limited generalizability due to the ambiguity of emotion perception and the diversity of data scenarios. To tackle this issue, we introduce UniEmoX, a cross-modal semantic-guided large-scale pretraining framework. Inspired by psychological research emphasizing the inseparability of the emotional exploration process from the interaction between individuals and their environment, UniEmoX integrates scene-centric and person-centric low-level image spatial structural information, aiming to derive more nuanced and discriminative emotional representations. By exploiting the similarity between paired and unpaired image-text samples, UniEmoX distills rich semantic knowledge from the CLIP model to enhance emotional embedding representations more effectively. To the best of our knowledge, this is the first large-scale pretraining framework that integrates psychological theories with contemporary contrastive learning and masked image modeling techniques for emotion analysis across diverse scenarios. Additionally, we develop a visual emotional dataset titled Emo8. Emo8 samples cover a range of domains, including cartoon, natural, realistic, science fiction and advertising cover styles, covering nearly all common emotional scenes. Comprehensive experiments conducted on six benchmark datasets across two downstream tasks validate the effectiveness of UniEmoX. The source code is available at https://github.com/chincharles/u-emo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。