构建首个百万级跨文化多语言菜系视觉问答基准
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
- 覆盖30种语言、9大语系,含超百万图文数据点
- 模型在正确位置上下文表现较好,但对地域特色菜系识别弱
- 适合研究跨文化理解、多语言视觉模型的学者使用
视觉语言模型(VLMs)普遍缺乏对非英语及少数文化背景知识的理解能力。为此,我们提出WorldCuisines,一个大规模多语言、多文化的视觉问答基准。该基准包含跨30种语言和方言、涵盖9个语言家族的视觉问答数据集,共超过100万条图文样本,是当前最大规模的跨文化视觉问答数据集。任务包括识别菜肴名称及其起源地。提供两种评估数据集(12,000 和 60,000 实例)以及一个100万实例的训练集。实验表明,尽管模型在正确地理位置上下文中表现更好,但在对抗性情境下仍难以准确预测特定区域菜系与语言。为支持后续研究,我们发布了包含标注食物条目与图像的知识库。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicultural, visually grounded language understanding. This benchmark includes a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects, spanning 9 language families and featuring over 1 million data points, making it the largest multicultural VQA benchmark to date. It includes tasks for identifying dish names and their origins. We provide evaluation datasets in two sizes (12k and 60k instances) alongside a training dataset (1 million instances). Our findings show that while VLMs perform better with correct location context, they struggle with adversarial contexts and predicting specific regional cuisines and languages. To support future research, we release a knowledge base with annotated food entries and images along with the VQA data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。