构建首个情感图像内容分析基准,评估视觉语言模型的情感理解与生成能力。
AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis
- 设计三任务综合评测框架:情绪理解、推理与引导生成。
- 发现23个模型普遍存在强度校准差、描述深度不足问题。
- 提出无需训练的分层提示方法,提升情绪表达准确性和丰富度。
视觉语言模型(VLMs)在感知任务中表现优异,但将感知、推理与生成整合于一体的全面情感图像内容分析(AICA)仍缺乏系统研究。为此,我们提出AICA-Bench,一个包含三大核心任务的综合性基准:情绪理解(EU)、情绪推理(ER)和情绪引导内容生成(EGCG)。我们评估了23个VLMs,发现其存在两大局限:情绪强度校准能力弱,以及开放性描述浅显。为解决此问题,我们提出无需训练的“有根基的情感树”(GAT)提示框架,结合视觉结构与分层推理。实验表明,GAT显著降低情绪强度误差,提升描述深度,为未来情感多模态理解与生成研究提供强有力基线。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, remains underexplored. To address this gap, we introduce AICA-Bench, a comprehensive benchmark with three core tasks: Emotion Understanding (EU), Emotion Reasoning (ER), and Emotion-Guided Content Generation (EGCG). We evaluate 23 VLMs and identify two major limitations: weak intensity calibration and shallow open-ended descriptions. To address these issues, we propose Grounded Affective Tree (GAT) Prompting, a training-free framework that combines visual scaffolding with hierarchical reasoning. Experiments show that GAT reduces intensity errors and improves descriptive depth, providing a strong baseline for future research on affective multimodal understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。