为波兰语设计的视觉语言评估基准,专测文化语境下的多模态理解。
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

- 基于波兰文化构建的单文化视觉问答数据集,语言与图像深度互动
- 含1117张图像和2366个标注问题对,挑战深层语义理解
- 适合研究跨文化视觉语言模型或波兰语NLP的学者使用
视觉语言模型在图像描述、视觉问答和图文生成等任务上表现强劲,但主要基于英语数据训练,难以处理文化相关的视觉理解,导致对区域特定含义、象征内容和上下文依赖视觉线索的误判。现有文化能力评估基准多依赖模板,仅关注表层识别,无法衡量文化情境下的深层语言与语用理解。我们提出PoVisLE,一个面向波兰语的单文化视觉语言评估基准,采用基于语境的评估范式,使语言在视觉背景中被解读。该数据集包含1117张图像和2366个手工标注的VQA对,提供受控且具有挑战性的资源,用于评估超越表层识别的文化化多模态理解能力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。