arXiv:2608.03358cs.CLcs.CV2026-08

构建跨文化视觉情绪理解基准,揭示大模型在文化感知上的短板。

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

论文配图:ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
图 1 · 摘自论文原文
  • 设计跨文化情绪理解任务,基于艺术作品标注文化特异性情绪
  • 16个大模型零样本测试准确率均低于50%,凸显挑战性
  • 引入检索增强框架,提升模型对文化语境的理解与解释能力

现有视觉情绪理解方法通常忽略情感认知的文化差异。本文提出文化条件下的视觉情绪理解任务,旨在预测图像在特定文化中的情绪感知并给出解释依据。现有基准存在个体标注不一致、缺乏多数共识文化标签及文化覆盖不平衡等问题。为此,我们构建了ArtECulture基准,包含6,792幅跨文化艺术作品,涵盖英语、汉语和阿拉伯语文化,东西方内容均衡。对16个开源与闭源多模态大模型(MLLMs)的零样本评估显示,最佳模型准确率不足50%。为解决此问题,我们提出一种检索增强的文化情绪理解框架,通过概念驱动的文化情绪知识库,在无需额外训练的前提下向MLLM注入显式文化知识,显著提升文化一致性情绪预测与可解释性生成。数据集与代码将公开发布。

原文摘要 · Abstract (English)

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.

跨文化理解视觉情绪多模态模型知识注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。