arXiv:2510.11178cs.CVcs.CY2025-10Conference of the …被引 5

测试视觉语言模型在跨文化语境下的理解能力,发现其易受语言表达变化影响。

BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models

  • 构建跨文化多模态问答基准,涵盖16个地区313个问题模板。
  • 模型在语言改写下性能下降明显,视觉信息提升有限且跨模态一致性低。
  • 适合研究文化敏感性、多模态对齐的AI开发者与评估者使用。

随着视觉语言模型(VLMs)在全球范围部署,其对文化情境知识的理解能力日益重要。然而,现有评估主要聚焦静态记忆或孤立视觉定位,无法检验VLM是否具备稳健且可迁移的文化理解能力。本文提出BLEnD-Vis,一个面向多模态、多文化的评估基准,用于检验VLM在不同语言表述和视觉模态下对日常文化知识的鲁棒性。基于BLEnD数据集,构造了313个文化相关的问题模板,覆盖16个地区,生成三种对齐的多项选择格式:(i) 仅文本基线(区域→实体),(ii) 反向文本变体(实体→区域),(iii) 带生成图像的VQA风格版本(ii)。最终包含4,916张图像和超过21,000个多项选择题实例,经人工标注验证。结果表明,当前VLM在文化知识上存在显著脆弱性,语言重述导致性能下降;尽管视觉线索常有帮助,但跨模态一致性低,尤其在低资源地区更明显。BLEnD-Vis为系统分析文化鲁棒性和多模态对齐提供了关键测试平台,揭示局限并指导更具文化适应性的模型发展。代码已开源于https://github.com/Social-AI-Studio/BLEnD-Vis。

原文摘要 · Abstract (English)

As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding. We introduce BLEnD-Vis, a multimodal, multicultural benchmark designed to evaluate the robustness of everyday cultural knowledge in VLMs across linguistic rephrasings and visual modalities. Building on the BLEnD dataset, BLEnD-Vis constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats: (i) a text-only baseline querying from Region $\rightarrow$ Entity, (ii) an inverted text-only variant (Entity $\rightarrow$ Region), and (iii) a VQA-style version of (ii) with generated images. The resulting benchmark comprises 4,916 images and over 21,000 multiple-choice questions (MCQ) instances, validated through human annotation. BLEnD-Vis reveals significant fragility in current VLM cultural knowledge; models exhibit performance drops under linguistic rephrasing. While visual cues often aid performance, low cross-modal consistency highlights the challenges of robustly integrating textual and visual understanding, particularly in lower-resource regions. BLEnD-Vis thus provides a crucial testbed for systematically analysing cultural robustness and multimodal grounding, exposing limitations and guiding the development of more culturally competent VLMs. Code is available at https://github.com/Social-AI-Studio/BLEnD-Vis.

多模态文化理解视觉问答评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。