MMArt为艺术理解提供多视角数据,突破单一描述局限。
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

- 构建四视角统一标注数据集:叙事、形式、情感、历史并行
- 实证表明各视角信息互补,历史视角在任务中不可替代
- 适合艺术理解、跨模态生成与模型评估研究者使用
当前视觉语言模型在艺术理解上仍停留在表面描述,缺乏形式分析、历史背景和情感特征的深入解读。现有数据集均为单一视角,无法同时提供叙事、形式、情感与历史视角。本文提出MMArt,包含74,234幅WikiArt作品,每幅画均配有四种独立标注视角及统一整合的描述,由专业模型或人工标注,并通过多重质量验证。互补性分析表明各视角编码独特信息;生成分析显示形式描述最保留学派风格,历史描述蕴含强情感信号;判别检索分析揭示任务不对称性:叙事描述在检索中表现最佳(R@1 = 44.0%),而形式描述虽在重建中最强,但在检索中几乎无区分力(R@1 = 7.8%)。留一分析进一步证实历史视角在两类任务中均最不可替代。结果证明单一视角不足以支持全面艺术理解,直接推动了MMArt多视角设计。数据与代码已开源。
原文摘要 · Abstract (English)
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task-asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave-one-out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi-perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。