让图像情绪分析更透明,能解释每种情绪的来源。
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
- 用视觉注意力提取关键区域和上下文线索
- 在三个数据集上比现有方法提升12.3%平均性能
- 适合需要可解释性的社交媒体情感分析场景
在线分享的图像对情绪和公众福祉有显著影响。理解图像引发的情绪对构建更健康可持续的数字社区至关重要,尤其在公共危机期间。本文研究视觉情绪诱发(VEE),即预测一张图像会引发哪些情绪。提出VAEER框架,一种可解释的多标签VEE方法,结合注意力启发的线索提取与知识引导推理。VAEER识别显著视觉焦点和上下文信号,将其与结构化情感知识对齐,并进行逐情绪推理,生成透明、情绪特异的解释。在三个异构基准上,包括社交图像和灾害相关照片,该方法达到领先性能,单情绪最高提升19%,平均比强基线(CNN和VLM)高出12.3%。研究结果表明,可解释的多标签情绪诱发是负责任视觉媒体分析和情感可持续在线生态系统的可扩展基础。
原文摘要 · Abstract (English)
Images shared online strongly influence emotions and public well-being. Understanding the emotions an image elicits is therefore vital for fostering healthier and more sustainable digital communities, especially during public crises. We study Visual Emotion Elicitation (VEE), predicting the set of emotions that an image evokes in viewers. We introduce VAEER, an interpretable multi-label VEE framework that combines attention-inspired cue extraction with knowledge-grounded reasoning. VAEER isolates salient visual foci and contextual signals, aligns them with structured affective knowledge, and performs per-emotion inference to yield transparent, emotion-specific rationales. Across three heterogeneous benchmarks, including social imagery and disaster-related photos, VAEER achieves state-of-the-art results with up to 19% per-emotion improvements and a 12.3% average gain over strong CNN and VLM baselines. Our findings highlight interpretable multi-label emotion elicitation as a scalable foundation for responsible visual media analysis and emotionally sustainable online ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。