构建28语言艺术图像描述数据集,推动跨文化视觉语言理解。
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
- 构建涵盖28种语言的20万条艺术图像描述,每图140条注释。
- 跨语言迁移在文化相近语言间效果更好,零样本学习表现显著。
- 适合研究多语言、跨文化视觉语言模型的研究者使用。
视觉与语言研究得益于COCO等基准数据集的推动,但传统数据集如COCO仅聚焦英语中的明确事实,ArtEmis引入主观情感,ArtELingo则初步拓展至中文和阿拉伯语。然而我们认为应进一步增强多语言覆盖。因此,我们提出ArtELingo-28,一个涵盖28种语言的视觉语言基准,包含约20万条标注(每张图像140条)。该数据集强调不同语言与文化背景下的观点多样性,挑战在于让机器为图像生成情感性描述。本文提供三种新设定下的基线结果:零样本、少样本和一对多零样本。实验表明,跨语言迁移在文化相关语言间更成功。数据与代码已公开于www.artelingo.org。
原文摘要 · Abstract (English)
Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjective emotions and ArtELingo introduced some multilinguality (Chinese and Arabic). However we believe there should be more multilinguality. Hence, we present ArtELingo-28, a vision-language benchmark that spans $\textbf{28}$ languages and encompasses approximately $\textbf{200,000}$ annotations ($\textbf{140}$ annotations per image). Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo-28 emphasizes diversity of opinions over languages and cultures. The challenge is to build machine learning systems that assign emotional captions to images. Baseline results will be presented for three novel conditions: Zero-Shot, Few-Shot and One-vs-All Zero-Shot. We find that cross-lingual transfer is more successful for culturally-related languages. Data and code are provided at www.artelingo.org.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。