arXiv:2509.16517cs.CVcs.AI2025-09EMNLP被引 13

构建跨文化视觉推理新基准,测试模型对东南亚文化图像的理解与定位能力。

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

  • 分两阶段评估:先选正确图像,再定位文化物件作为证据
  • 涵盖7国138种文化物件,3178个问题由人工精心设计
  • 揭示当前模型在文化推理与空间定位间的显著差距

多模态视觉语言模型在需结合视觉与文本理解的任务中取得显著进展,尤其在文化理解方面。然而现有文化数据集常缺乏深层文化推理能力,且多数文化代表性不足。本文提出「Seeing Culture Benchmark(SCB)」,聚焦文化推理,采用双阶段方法:第一阶段为多选视觉问答(VQA),要求模型从三类选项中选出正确图像——同国、异国或混合类别;第二阶段仅在第一阶段正确后触发,要求模型分割出支持推理的文化物件。所有选项均来自同一类别。SCB包含1,065张图像,覆盖东南亚七国的138种文化物件,分属五个类别,配套3,178个问题,其中1,093个为人工精标。评估显示,当前多种模型在跨模态文化推理中存在复杂性,且视觉推理与空间定位能力不匹配。该基准有助于识别缺陷,推动文化推理研究发展。

原文摘要 · Abstract (English)

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures. In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual options in the first stage are systematically organized into three types: those originating from the same country, those from different countries, or a mixed group. Notably, all options are derived from a singular category for each type. Progression to the second stage occurs only after a correct visual option is chosen. The SCB benchmark comprises 1,065 images that capture 138 cultural artifacts across five categories from seven Southeast Asia countries, whose diverse cultures are often overlooked, accompanied by 3,178 questions, of which 1,093 are unique and meticulously curated by human annotators. Our evaluation of various VLMs reveals the complexities involved in cross-modal cultural reasoning and highlights the disparity between visual reasoning and spatial grounding in culturally nuanced scenarios. The SCB serves as a crucial benchmark for identifying these shortcomings, thereby guiding future developments in the field of cultural reasoning. https://github.com/buraksatar/SeeingCulture

视觉推理文化理解多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。