arXiv:2602.01173cs.CV2026-02被引 3

构建最大图像情感数据集,提升机器对图像引发情绪的多维度理解能力。

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

  • 基于125K图像生成120万问答对,覆盖5维情绪分析任务。
  • 提出EEmo-Logic模型,在跨域数据上实现情绪问答与细粒度评估双优表现。
  • 适合研究人机共情、多模态情感计算的学者使用。

理解图像引发情绪的多维属性与强度细微差别,对推动机器共情及丰富人机交互应用至关重要。现有模型仍局限于粗粒度情绪感知或推理能力不足。为此,我们推出迄今为止最大的图像诱发情绪理解数据集EEmoDB,涵盖5个分析维度和5类任务,支持全面解读。具体包括:通过自动化生成从12.5万张图像中获得120万条问答对(EEmoDB-QA),以及从2.5万张图像中人工精标3.6万条细粒度评估数据(EEmoDB-Assess)。此外,我们提出EEmo-Logic,一种通过指令微调和定制化组相对偏好优化(GRPO)训练的统一多模态大模型,采用新颖奖励设计。大量实验表明,EEmo-Logic在领域内和跨域数据集上均表现稳健,情绪问答与细粒度评估性能优异。数据集与代码已公开于https://github.com/workerred/EEmo-Logic。

原文摘要 · Abstract (English)

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still limited to coarse-grained emotion perception or deficient reasoning capabilities. To bridge this gap, we introduce EEmoDB, the largest image-evoked emotion understanding dataset to date. It features $5$ analysis dimensions spanning $5$ distinct task categories, facilitating comprehensive interpretation. Specifically, we compile $1.2M$ question-answering (QA) pairs (EEmoDB-QA) from $125K$ images via automated generation, alongside a $36K$ dataset (EEmoDB-Assess) curated from $25K$ images for fine-grained assessment. Furthermore, we propose EEmo-Logic, an all-in-one multimodal large language model (MLLM) developed via instruction fine-tuning and task-customized group relative preference optimization (GRPO) with novel reward design. Extensive experiments demonstrate that EEmo-Logic achieves robust performance in in-domain and cross-domain datasets, excelling in emotion QA and fine-grained assessment. The dataset and code are available at https://github.com/workerred/EEmo-Logic.

情绪理解多模态大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。