arXiv:2503.21910cs.CLcs.AI2025-03Conference of the …被引 13

评测视觉语言模型在四种阿拉伯方言中的理解能力,发现现有模型表现不佳。

JEEM: Vision-Language Understanding in Four Arabic Dialects

  • 构建四国方言视觉理解基准JEEM,涵盖图像描述与视觉问答任务。
  • 五款开源阿拉伯VLM均表现较差,GPT-4V最优但方言差异下能力不稳定。
  • 强调需更包容的模型和文化多样性的评估体系,适合多语种AI研究者。

我们提出JEEM,一个针对约旦、阿联酋、埃及和摩洛哥四个阿拉伯语国家的视觉语言模型(VLM)评估基准。该数据集包含图像描述和视觉问答任务,内容富含文化特色且地域多样。旨在评估VLM在不同方言间的泛化能力以及对视觉情境中文化元素的准确理解。对五款主流开源阿拉伯VLM及GPT-4V的评估显示,阿拉伯VLM普遍表现不佳,既在视觉理解上受限,又难以生成符合方言特征的文本。尽管GPT-4V在对比中表现最佳,其语言能力在不同方言间存在差异,视觉理解能力也相对滞后。这凸显了开发更具包容性的模型的必要性,以及采用文化多样性评估范式的价值。

原文摘要 · Abstract (English)

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image captioning and visual question answering, and features culturally rich and regionally diverse content. This dataset aims to assess the ability of VLMs to generalize across dialects and accurately interpret cultural elements in visual contexts. In an evaluation of five prominent open-source Arabic VLMs and GPT-4V, we find that the Arabic VLMs consistently underperform, struggling with both visual understanding and dialect-specific generation. While GPT-4V ranks best in this comparison, the model's linguistic competence varies across dialects, and its visual understanding capabilities lag behind. This underscores the need for more inclusive models and the value of culturally-diverse evaluation paradigms.

视觉语言多语言阿拉伯语评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。