arXiv:2603.21165cs.CLcs.CV2026-03中稿 · EMNLP被引 1

构建多语言方言文化评估基准,测试模型对孟加拉文化的理解能力。

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

  • 基于9个领域1152张人工标注图,构建涵盖5种方言的多语言文化评估集
  • 模型在方言变化下表现下降,尤其在图文生成任务中,文化知识缺失是主要瓶颈
  • 适合研究跨语言文化理解、多模态模型泛化能力的研究者使用

孟加拉文化通过地域、方言、历史、饮食、政治、媒体和日常视觉生活丰富呈现,但在多模态评估中仍被忽视。为此,我们提出 BanglaVerse,一个以文化为基础的多语言视觉-语言模型(VLMs)评估基准,用于衡量在历史上相关语言和区域方言中对孟加拉文化的理解。该基准包含1,152张人工标注图像,覆盖九个领域,支持视觉问答与图文生成任务,并扩展至四种语言和五种孟加拉语方言,共生成约32.2万条数据。实验表明,仅以标准孟加拉语评估会高估模型真实能力:在方言变化下性能下降,尤其是图文生成任务;虽然印地语、乌尔都语等历史关联语言保留部分文化含义,但在结构化推理上仍较弱。跨领域分析显示,主要瓶颈是文化知识缺失,而非视觉定位本身,尤其在知识密集型类别中。这些发现使 BanglaVerse 成为衡量语言变异下文化根基多模态理解能力的更真实测试平台。

原文摘要 · Abstract (English)

Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision-language models (VLMs) on Bengali culture across historically linked languages and regional dialects. Built from 1,152 manually curated images across nine domains, the benchmark supports visual question answering and captioning, and is expanded into four languages and five Bangla dialects, yielding ~32.2K artifacts. Our experiments show that evaluating only standard Bangla overestimates true model capability: performance drops under dialectal variation, especially for caption generation, while historically linked languages such as Hindi and Urdu retain some cultural meaning but remain weaker for structured reasoning. Across domains, the main bottleneck is missing cultural knowledge rather than visual grounding alone, with knowledge-intensive categories. These findings position BanglaVerse as a more realistic test bed for measuring culturally grounded multimodal understanding under linguistic variation.

文化理解多语言视觉语言模型方言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。