让AI看画作时只关注真正影响情绪的细节,避免罗列无关元素。
Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language Models

- 基于预设形式属性构建可解释推理框架,筛选真正影响情绪的视觉线索。
- 在EmoArt数据集上提升情绪、唤醒度预测准确率,解释文本更简洁。
- 适合需要可解释艺术情感分析的研究者与跨模态生成应用。
多模态大模型能生成流畅的艺术品情绪解释,但常出现属性泛滥问题:过度列举可见的形式属性,却未识别真正支撑情感判断的关键线索。为此,本文将艺术品情绪理解建模为属性基础的选择性推理(AGSR),其中预定义的形式属性作为证据单元,仅情绪相关属性应进入最终解释。为使该问题可度量,本文在原始132,664幅作品的EmoArt数据集基础上,新增1,400幅由15名艺术训练标注者标注的显著性扩展数据,提供实例级监督以区分仅存在与情绪显著的属性。我们提出FAB-G(形式属性瓶颈引导推理)——一种监督式多智能体框架,先预测属性显著性,再约束下游情绪分析仅使用保留线索。实验表明,FAB-G在情绪、唤醒度和效价预测上持续提升,与人工标记显著属性在Dice和Tversky指标上达成更强一致,且生成解释显著更紧凑。跨数据集评估进一步表明,属性基础的显著性选择可迁移至EmoArt分布之外,同时揭示特定属性的边界案例。数据集与项目页见https://zhiliangzhang.github.io/EmoArt-130k/
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) can produce fluent artwork emotion explanations, but they often suffer from attribute flooding: they enumerate many visible formal attributes without identifying which cues actually support the affective judgment. We therefore formulate artwork emotion understanding as Attribute-Grounded Selective Reasoning (AGSR), where predefined formal attributes serve as evidence units and only emotionally operative attributes should enter the final interpretation. To make this problem measurable, we extend EmoArt, originally introduced at ACM MM 2025 as a 132,664-artwork resource with content, formal-attribute, valence-arousal, and emotion annotations, by adding a 1,400-artwork human salience extension annotated by 15 art-trained annotators. This extension provides instance-level supervision for distinguishing attributes that are merely present from those that are emotionally salient. We further propose FAB-G (Formal-Attribute Bottleneck-Guided reasoning), a supervised multi-agent framework that first predicts attribute-level salience and then constrains downstream emotional analysis to the retained cues. Experiments show that FAB-G yields consistent gains in emotion, arousal, and valence prediction, achieves stronger agreement with human-marked salient attributes under Dice and Tversky metrics, and produces substantially more compact final explanations than prompting-based baselines. Cross-dataset evaluation further suggests that attribute-grounded salience selection transfers beyond the source distribution of EmoArt, while also revealing attribute-specific boundary cases. The dataset and project page are available at https://zhiliangzhang.github.io/EmoArt-130k/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。